My thoughts on the foreign policy AI evaluation gap

A recent paper published on Arxiv, 'The Foreign Policy AI Evaluation Gap', highlights the complexity of grading the use of AI when it comes to statecraft, which they define as "the purposive, institutionally mediated conduct through which a political actor seeks to shape its external environment." So this covers the whole toolkit that a state can use to act on the world, which includes everything from imposing sanctions, managing a crisis, using force, to drafting agreements, and also negotiating.

In the paper, the authors claim that rather than asking whether AI is good for foreign policy in general, which is impossible to answer due to a lot of moving variables in real life statecraft, a five stage workflow should be tested at every turn and see whether AI usage for that task only has been successful or not when compared to just a human doing the work. They break this workflow into Research, Analyse, Strategise, Execute, and Monitor, with each one having its own 'Task', 'Output', 'Evaluation', and 'Primary TAIG (technical AI governance) capacities'. From my experience as a former diplomat of an EU member state, these five stages reflect how diplomacy, though not statecraft as a whole, produces an outcome.

This framework for foreign-policy AI systems as suggested by Pozniak and Sania works well in a lot of cases I worked on.

Let's take for example the preparations for an EU Foreign Affairs Council. Before an FAC session, we would map who wants what and where the constraints sit (Research), draft lines to take, and have our Permanent Representation give us feedback based on what they heard in Brussels. Positions rarely shift once the session starts, since most lines to take are settled in advance, but in other cases such as heated negotiations, a state's position can move, which is where you need to work out whether there is a real shift in language (Analyse). From there, usually the member state pushes its own line, or aligns with a group of states who have a common aim (Strategise).

The last two stages do not exactly fit in the FAC case, as foreign policy is a member state competence. That said, the High Representative of the Union for Foreign Affairs and Security Policy issues common positions and statements once these have been negotiated between all member states (Execute). Monitor, meanwhile, can refer to checking whether a sanctions listing, approved by member states at the Council, still holds up at renewal, or whether new information means a case needs to be flagged.

This framework runs into what the paper calls 'contested ground truth' - information this kind of AI system cannot capture such as the feedback we got from the PermRep in Brussels, for example, which could have been unpublished official statements from other member states, but also comments in passing or one-to-one conversations. An AI system doing the 'Research' part from public statements alone would have less data than officials would have.

The framework has the right shape, and the paper makes a strong case for it. I look forward to seeing someone build it. Andon Labs for diplomacy anyone? Minus the 'no humans in the loop' part, of course.