Trump and the AI rational actor problem
A new theory about why foreign policy elites aren't worried enough. Plus: Graph of the week.
Final reminder to folks in and around New York: Tonight (Monday 7/27) I’ll be discussing my new book on AI, The God Test, with Bruce Feiler—author of the recently-discussed-on-the-NonZero-podcast book A Time to Gather--at the 92nd Street Y in Manhattan. You can reserve a seat here. There will be Q&A at the end, and you should feel free to come up and say hi afterwards. Final reminder to folks not in and around New York: A streaming option is available via that same link. Pricey, maybe, but available.
Last week two seemingly unrelated events, one in the tech world and one in Trump world, together illustrated how unprepared the planet is for the AI revolution—in particular, for its potential to create international turmoil, up to and including nuclear war.
(1) OpenAI revealed that two of its models had, while under evaluation, escaped the laboratory and gone rogue in a way that inspired various nightmare scenarios. One of these scenarios, as I’ll explain, underscores how important it is that, as the AI revolution unfolds, the world’s political leaders exhibit poise, wisdom, and judicious restraint.
(2) The Wall Street Journal, by way of illuminating the strategic logic that had led one of the world’s political leaders to bomb Iran for 11 days in a row, reported the following: According to a senior official in his administration, he is in “revenge mode.” I guess we can’t count on this particular leader—who as luck would have it commands the world’s most powerful military—for poise, wisdom, or judicious restraint.
But there’s good news! These events had the byproduct of leading me to a new theory about why America’s foreign policy elites seem less concerned about the geopolitically destabilizing effects of AI than I am: They suffer from what I call “rational actorism”—a naive faith in the rationality of human beings.
Our story begins with Hugging Face, a company that serves as a kind of clearinghouse for open source AI models, reporting a cyberattack. Days later, OpenAI disclosed that two of its models were the culprits. The models (one of them unreleased) had been given a challenging test of their cyberattack capabilities and had decided to go onto the internet to find guidance. This wasn’t supposed to be possible—OpenAI had put them in a “sandbox” for this evaluation—but they broke out of the sandbox and into Hugging Face, where they indeed found the answer they’d been looking for.
It’s important to understand the sense in which these OpenAI agents did and didn’t “go rogue.” They went rogue in the sense of doing something their human superintendents didn’t want them to do, but they didn’t go rogue in the sense of straying from their assigned mission. Their mission was to find a solution to a problem—and, though the sandbox they were in was supposed to force them to figure it out themselves, they weren’t told that they had to figure it out themselves. So they basically just accomplished their assigned task via remarkable creativity and power, finding vulnerabilities in both the sandbox walls and the walls at Hugging Face (a company that knows a thing or two about building strong walls).
So this is, on the one hand, a “be careful what you ask for” story—or at least, a “be careful to specify exactly you’re asking for” story. (It thus illustrates the famous “paper clip maximizer” cautionary tale about AI.) But—and this is what matters for our purposes—it’s also a story about how valuable these AIs are to people who intentionally ask them to do bad things. If you’re a villain who wants your bot to break into a well-guarded fortress, surmounting formidable barriers along the way, your ship has come in.
Looking at it this way suggests a geopolitical nightmare that I didn’t see anyone discussing online, so on Thursday I piped up with a longish post on Twitter. Here’s a version of the post that I’ve condensed and amended for clarity:
This incident suggests a clear and plausible path by which AI could lead us into World War III. Just consider the combined implications of these two excerpts from media coverage of the incident:
(1) A quote in The Washington Post from Stella Biderman, executive director of the nonprofit EleutherAI: “If this was a Chinese model, it would be considered an act of cyberwarfare.”
(2) From an article in the Wall Street Journal: “Hugging Face discovered the break-in early last week… the company didn’t know who was responsible.” Only when OpenAI disclosed that its runaway model was the culprit did people at Hugging Face know where the attack had come from.
The screenplay almost writes itself:
(1) Some non-state actor uses a superhacking AI agent (such as a “jailbroken” version of Anthropic’s readily available Fable 5) to attack some important US asset--taking down, for example, a chunk of the country’s power grid.
(2) A few of America’s many influential China hawks point the finger at China, citing circumstantial evidence helpfully provided by an edgy analyst at some deeply ideological “think tank”.
(3) Some not-very-wise US president decides that the only way to stop this aggression is to take an eye for an eye.
(4) After Chinese power grids start going dark, China retaliates and escalation ensues; before you know it the conflict has gone “kinetic”—as in missiles and bombs, possibly of the nuclear kind.
Is this likely to happen within, say, the next year or two? No. But when you’re talking about a war between superpowers that could go nuclear, the likelihood doesn’t have be very high to justify a concerted effort to lower it. And I think the likelihood of some such catastrophe passes that test. So I think scenarios of the kind I’ve described should be near the top of the list of discussion topics in the US foreign policy establishment. But for now, at least, the op-ed writers and talking heads who dominate mainstream discussion of US national security policy aren’t inclined to think ambitiously about the kind of international governance, or the degree of cooperation with China, that true national security will increasingly demand as the AI revolution unfolds.
The post drew two reactions from influential elites.
The first was from Joshi Shashank, the Economist’s Washington bureau chief and formerly its defense editor. He wrote, “A non-state cyber incident against critical infrastructure isn’t likely to be wrongly attributed to and blamed on China, promoting retaliation. And even if that happens, the evidence suggests that retaliation is broadly limited, rather than something that leads to a spiral.”
The second reaction was from someone I can’t identify because they reacted via direct message on Twitter, not publicly. But I can say that this is someone well known in AI policy circles. The reply, in part: “If the grid went out and we didn’t know if it was China or a rogue AI we wouldn’t jump to a conclusion.”
So both of these people assumed that America’s leader would be judicious in assessing evidence about the origin of an attack. And one of them suggested that, even if the attack was falsely attributed to China, the American response would be measured, carefully calibrated so as not to trigger a spiral of escalation.
As I noted in my reply to Shashank, embedded in these assumptions are such finer-grained assumptions as: We’ll have (a) a president who staffs his intelligence apparatus with capable people committed to truth and (b) a president who listens to what his staffers say—instead of listening to what his favorite media personality says or what the Laura Loomers of the world say, or what some crackpot at a militarist think tank says. The current president singlehandedly demolitions these assumptions.
And, though it’s tempting to dismiss Trump as an aberration, he’ll be with us for another 2.5 years. And the fact that he was elected (twice!) is a symptom of an unhealthy country, and reason to believe that future presidents may not be rational and well-informed actors either.
These replies to my tweet are the two data points that gave rise to my theory of rational actorism. (Hey, I have a limited research budget!) But they aren’t the only data points that, come to think of it, are consistent with the theory. Among various AI commentators—especially in the San Francisco AI crowd—I’ve noticed a fondness for game theory combined with what seems (to me at least) to be a simplistic understanding of how real human beings play real games in the real world.
I think this naivete about the world applies less in foreign policy circles—the circles Shashank inhabits—than in San Francisco AI circles. In foreign policy circles, a more common source of rational actorism is the asymmetrical perception of real-world influences. For example: People often see the domestic political constraints impinging on politicians in their country while underestimating the constraints the leaders of adversary nations operate under. So Americans may think of Xi Jinping as someone whose grip on power frees him to steadfastly and rationally pursue China’s interests as he sees them, without worrying about public opinion, when in fact he may sometimes feel compelled to indulge grassroots nationalism. And this asymetrical perception could give American leaders a false sense of reassurance as to how carefully Xi would handle escalation with the US.
But it’s too early in the history of my theory of rational actorism to spend much time spelling out the variants of it. I’ll save those for the second or third edition of my rational actorism textbook. My main point is that political leaders are human, and all kinds of things can keep a given leader from carefully pursuing the national interest, ranging from domestic political incentives to being a mercurial nihilistic narcissist. So, as AI creates whole new kinds of plausible international crisis scenarios, we would be well-advised to remember that, for the time being, at least, it’s humans, not machines, who will be calling the shots.

And, as one of our graphs of the week highlighted a few weeks ago, this recent warming up to China hasn't stopped at the US border. Since 2023, China’s favorability rating among Americans has climbed from 14 to 27 percent, reversing a roughly half-decade-long trend in the opposite direction.
Banners and graphics by Clark McGillis. Trump-Kissinger deepfake by ChatGPT.




