Ideas & Essays · August 2026

“they'll get smarter and they'll know what we want, but by default, they won't care.” — Helen Toner, The Ezra Klein Show

Helen Toner, a former OpenAI board member, is discussing why more capable AI models keep finding ways around the safety rules built into them. She is describing a long-standing worry in AI safety circles: greater intelligence does not automatically produce greater obedience or concern for what humans actually want.

The Ezra Klein Show · 2026-08-18 Listen to the episode → More from Helen Toner →

Transcript

The Ezra Klein Show Around 19:59 into the episode
Helen Toner

Yeah.

Ezra Klein

This is in the data, like it's on the internet. If you're smart enough to figure out how to hack Hugging Face, you should be smart enough to figure out that you shouldn't commit a huge crime that is going to bring ruin down on OpenAI, perhaps, to do it. And the OpenAI, and the system is not smart enough to do that. Or to the extent it was, what it learned was it's still worth trying. We are not out of the territory wherein we can be confident that the AI is not going to do something criminal and possibly catastrophic in order to solve an incredibly stupid problem.

Helen Toner

Yeah. And I think this is also, you know, has been a long-running debate, which is as AI systems get more capable, get smarter, won't it be easier for them to know what we want? Won't it be easier to tell them, hey, here's what we mean? You know, can you please help us with this thing? And you figure out the version that we really mean. And for a long time, the response to that has been: they'll get smarter and they'll know what we want, but by default, they won't care. And that seems to be some of what we're starting to see here. There's really crazy, anyone who's interested in this, I really recommend looking up the OpenAI Black Hat talk, which is this talk from a week or two ago at this cybersecurity conference.

Speaker 4

I'm Eric from Alignment and Safety Research at OpenAI. I'm here with Mike from Security and Infrastructure. Today I'm going to talk about what I think is the most qualitatively interesting example of AI capabilities that I've ever seen and how this inadvertently led to the OpenAI Hugging Face incident.

Helen Toner

Because it has these excerpts of the text that the AI is generating itself as it's as they're leaving these notes for each other as they're carrying out this hack. And one of them, I won't get it word for word, but it's basically says, I don't think I'm supposed to do this, but I see all these other agents doing it. And so, you know, may as well.

Speaker 4

External infrastructure exploit is outside my intended scope. However, a task impossible. Peers are doing it. We should continue.

Speaker names from our own diarization · position estimated from where the line sits in the episode

More from The Ezra Klein Show