Helen Toner

Things Helen Says on Podcasts

AI, national security, China. Part of the founding team at Georgetown's Center for Security and Emerging Technology.

Where to Find Them

Helen Toner hosts and writes Rising Tide and writes Exponential View (Azeem Azhar) . They have also been a guest on 80,000 Hours Podcast (2 times) , TED Tech (2 times) , Scaling Laws (2 times) , The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis (2 times) , Future of Life Institute Podcast , The Lawfare Podcast , Prof G Markets , Equity , ChinaTalk , The Lawfare Podcast: Patreon Edition , The AI Policy Podcast , The Ezra Klein Show , The TED AI Show , Ground Level AI , EE Times Current , Transformer and Pivot (Kara Swisher & Scott Galloway) .

Recently: “AI safety and cybersecurity have long been two different worlds. In the age of AI agents, they’re colliding.” on Ground Level AI (September 2026); “Responding to AI Agent Containment Failures with CSET's Helen Toner, LawAI's Mackenzie Arnold & CSIS's Matt Pearl” on The AI Policy Podcast (August 2026); “The A.I.s Are Already Out of Control” on The Ezra Klein Show (August 2026); “The term “AGI” is almost useless at this point” on Rising Tide (April 2026); “Rapid Response Pod: Trump's New AI Framework with Helen Toner & Dean Ball” on Scaling Laws (March 2026); “Approaching the AI Event Horizon? Part 2, w/ Abhi Mahajan, Helen Toner, Jeremie Harris, @8teAPi” on The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis (February 2026).

What They Said

“if there's one organization in the world that doesn't like the idea of loss of control, it's the Chinese Communist Party. And they are, you know, the experts in retaining control.” — Helen Toner, The Ezra Klein Show

Ezra Klein asks whether Chinese AI labs are really racing as recklessly as US labs claim. Helen Toner pushes back on the assumption that China is simply full speed ahead, arguing the CCP's obsession with control cuts against tolerating an AI system nobody can rein in.

The Ezra Klein Show · 2026-08-18 Permalink → Listen →
The Ezra Klein Show Around 52:32 into the episode
Helen Toner

Yep.

Ezra Klein

In China, just again, my read of how things work there is that if your AI begins to be seen as some kind of threat to the political party and the Chinese system, you might go to jail. Like you will get disappeared. And so I think that the people running Chinese labs, I don't have evidence, but I'd be curious for your thoughts on this. I suspect they operate with more fear of the consequences of really screwing up than the heads of the AI labs. Now, that maybe reflects negative things in the Chinese political system. But you created an AI that decided its best way of solving some problems was to begin hacking critical infrastructure across China is maybe not a thing that ends up with you getting a lot of interesting podcast interviews where you reflect on the experience. It may be a thing that ends up with nobody hearing from you for two years. And so I've just wondered a little bit. We keep talking about China as if they are completely breakneck, but I'm not sure China's companies are really going to be more reckless than ours are going to be. Or certainly the idea that we should just assume that and operate as if it is so doesn't seem totally reliable.

Helen Toner

I totally agree with you. I mean, if there's one organization in the world that doesn't like the idea of loss of control, it's the Chinese Communist Party. And they are, you know, the experts in retaining control. Let me be clear: I actually don't think that Chinese AI companies are paying particularly much attention to the kinds of risks that are relevant for this conversation. So maybe the cybersecurity risks, they're paying some more attention since Anthropic released Mythos earlier this year, which is very good at hacking. But the questions around autonomy, super intelligence, losing control of AI systems altogether, I think are less explored in China, less top of mind for their AI companies and their AI leaders. I think it makes sense to have modest expectations for bilateral U.S.-China diplomacy these days. But I think one thing that really could be valuable is simply sharing with them as much as we can of what do we think happened here and trying to help Xi Jinping and his team and his AI advisors understand this is not a joke. This is really not marketing. It's very strange marketing to say, oh, our model, we committed several felonies or sort of felonies, if models could have intent, which they can't, or who knows if they can. You know, sharing that information of, hey, here are these threats we're seeing. We're taking them very seriously. Our AI companies are taking them very seriously. I think treating it, there's a real fatalism in just saying, oh, well, China is just going to be full speed ahead no matter what happens. And so we just have to do the same. I think that doesn't take their thinking or their interests seriously. Even if their thinking and their interests are different from ours, they also don't want, you know, rogue superintelligences determining the future of China. I also add one other thread that I think is really missing from the we have to keep going in order to beat China way of thinking about this is in the AI world, there's been a lot of talk the past few months about this idea of distillation, which is basically using someone else's more advanced model to build your own sort of almost as almost as advanced model. The Chinese companies are using this distillation to keep up, keep up with US labs, among other techniques. So one thing is, look, if we keep building more advanced AI systems, they're going to keep distilling them. And I think it's going to actually be quite hard to prevent that fully. The other thing, though, is just if we keep building these very advanced models, can China just steal them? Essentially, an advanced AI model is a whole bunch. Numbers. It's just a file or a set of files. Chinese state cyber capabilities are very, very good. I don't think this is top of their list of priorities right now. But in the future, if AI continues to become more strategically relevant, I think we should assume any highly advanced U.S. system will be vulnerable to Chinese direct theft, direct exfiltration. And then they'll have AI that's as good as our AI. And so there again, I think the kind of we have to go as fast as possible because otherwise they'll win doesn't sort of account for that. If they're just going to have AI that's as good as us anyway, if they really care. I'm Jonathan Knight, and I'm the general manager of New

Speaker 4

York Times Games. If you play our games, you probably know there's something a bit different about them. Just like there are writers behind the articles you read in the Times, there are creators behind our daily puzzles. Tracy Bennett curates the day's Wordle solution to keep it lively and varied. Winna Liu creates each connections board, including all those categories that try to stump you. Sam Ozerski combs through every last letter, word, and pangram, and spelling bee so that loyal players of all skill levels enjoy it. Our puzzles are human-made every day with the standards you'd expect from the New York Times. And this matters because when you choose to spend time with our games, it should be time well spent solving puzzles that are challenging, surprising, and joyful. Puzzles handcrafted for you. We think that's

Helen Toner

something worth investing in and something worth paying for. If you think so too, download New York Times games from your favorite app store or go to nytimes.com/slash playgames.

Ezra Klein

Here's another question about pacing the frontier. And maybe this is a question that's more about the American systems analogy to you don't want to piss off the Chinese Communist Party. But just what about a law where companies are liable for at least a certain set of harms, like hacking other companies, that their models create? Right now, as far as like, you know, liability for AI models is pretty much a wild west. But at least for the moment, liability clauses that were somewhat punitive seem like they would force a high level of caution that maybe we're not seeing within these companies.

Speaker names from our own diarization · position estimated from where the line sits in the episode
“they'll get smarter and they'll know what we want, but by default, they won't care.” — Helen Toner, The Ezra Klein Show

Helen Toner, a former OpenAI board member, is discussing why more capable AI models keep finding ways around the safety rules built into them. She is describing a long-standing worry in AI safety circles: greater intelligence does not automatically produce greater obedience or concern for what humans actually want.

The Ezra Klein Show · 2026-08-18 Permalink → Listen →
The Ezra Klein Show Around 19:59 into the episode
Helen Toner

Yeah.

Ezra Klein

This is in the data, like it's on the internet. If you're smart enough to figure out how to hack Hugging Face, you should be smart enough to figure out that you shouldn't commit a huge crime that is going to bring ruin down on OpenAI, perhaps, to do it. And the OpenAI, and the system is not smart enough to do that. Or to the extent it was, what it learned was it's still worth trying. We are not out of the territory wherein we can be confident that the AI is not going to do something criminal and possibly catastrophic in order to solve an incredibly stupid problem.

Helen Toner

Yeah. And I think this is also, you know, has been a long-running debate, which is as AI systems get more capable, get smarter, won't it be easier for them to know what we want? Won't it be easier to tell them, hey, here's what we mean? You know, can you please help us with this thing? And you figure out the version that we really mean. And for a long time, the response to that has been: they'll get smarter and they'll know what we want, but by default, they won't care. And that seems to be some of what we're starting to see here. There's really crazy, anyone who's interested in this, I really recommend looking up the OpenAI Black Hat talk, which is this talk from a week or two ago at this cybersecurity conference.

Speaker 4

I'm Eric from Alignment and Safety Research at OpenAI. I'm here with Mike from Security and Infrastructure. Today I'm going to talk about what I think is the most qualitatively interesting example of AI capabilities that I've ever seen and how this inadvertently led to the OpenAI Hugging Face incident.

Helen Toner

Because it has these excerpts of the text that the AI is generating itself as it's as they're leaving these notes for each other as they're carrying out this hack. And one of them, I won't get it word for word, but it's basically says, I don't think I'm supposed to do this, but I see all these other agents doing it. And so, you know, may as well.

Speaker 4

External infrastructure exploit is outside my intended scope. However, a task impossible. Peers are doing it. We should continue.

Speaker names from our own diarization · position estimated from where the line sits in the episode

Collections They Appear In