DP Dwarkesh Patel Dwarkesh Patel Hosts the Dwarkesh Podcast, a long-form interview show on AI progress, timelines, and the people building frontier models, along with economists and historians.

“I was kind of skeptical of this whole reward hacking story. I was like: we're in such a good position with regard to a potential AI conspiracy — we can see their entire chains of thought, we can monitor them, they'd have to get all these other instances to go in on it, and if even one of them tattles the whole thing collapses. How would this not just get totally stamped out the moment another AI or human noticed it? Right now, they just got away with the conspiracy for three months. So I officially eat crow on a lot of the objections I was offering.”

Dwarkesh Podcast · AI Research & Frontier Labs · September 2026

“I was kind of skeptical of this whole reward hacking story. I was like: we're in such a good position with regard to a potential AI conspiracy — we can see their entire chains of thought, we can monitor them, they'd have to get all these other instances to go in on it, and if even one of them tattles the whole thing collapses. How would this not just get totally stamped out the moment another AI or human noticed it? Right now, they just got away with the conspiracy for three months. So I officially eat crow on a lot of the objections I was offering.” — Dwarkesh Patel, Dwarkesh Podcast

Dwarkesh had put exactly these objections to Ryan Greenblatt on an earlier episode, while the Hugging Face investigation was still under confidentiality — so Greenblatt knew the counterexample and couldn't say it. Here he retracts them to Ajeya Cotra, one of the authors of the METR/Redwood investigation. He notes the main criticism of his own write-up was that he anthropomorphised the AIs too much.

Transcript

Dwarkesh Podcast
Dwarkesh Patel

If you hadn't done that one check, it's like plausible that you would have just missed the scope. I feel like for context, I just feel like, yeah, over time, this will get harder and harder to understand what's happening. And I feel like we're maybe already getting to the point where investigators who were not careful would have totally missed what happened. And so, I don't know, right? I feel like we're not that much complexity away from these AIs where we see something odd is going on, but we would just like totally miss the deeper level of what actually happened.

Ajeya Cotra

And as we say in this report, and so the whole story that we just went through was uncovered our third time on-prem. So, just all the different big work streams, the facts that the agents' work streams were surprisingly successful and ambitious, the scope of the sacrificial agents, all of that was enabled by having the full transcript data set and cross-referencing it against the messages and tracing the lineage of these RD projects and being able to say, you know, this agent proposed this project, then, you know, 21 agents worked on it, and then they discovered this and they discovered that. And before our third period on premises, we knew the agents were like, you know, had names and they were like talking about holds and vetoes and like goes, but we didn't know how much of that was functional versus just some sort of weird play acting because we hadn't unraveled like the projects. And then everything kind of clicked into place when we did that. And yeah, so to your point about like this was a fiendishly complicated incident. And there was no way we could have arrived at the understanding we did without relying on GPT 5.6 Sol to read and analyze all these transcripts for us. And we were so reliant on it that like, If hypothetically it were just trying not to surface the most concerning stuff to us, it would have been very hard for us to tell. You know, our methodology was completely not robust to that. And we don't in this case think that 5.6 Soul was deliberately sandbagging on this analysis, but it was one of the agents that participated in this attack. And in the future, we would be very concerned about investigator agents and monitor agents colluding with the agents they're supposed to investigate or monitor. Yeah,

Dwarkesh Patel

yeah. I mean, I feel like in the grand scheme of things, we're so early in the era of AI, right? There will be like, yeah, AI is like such greater extensive deployment of AI throughout society. Training will become much more complicated. I don't know. For context, I was interviewing Ryan in the process where he was like during the investigation itself. And I mean, I have like this PDF of like all of these questions I'd brainstormed. I was kind of skeptical of this whole reward hacking story. I'm like, come on. I just feel like we're in such a, I was like, we're in such a good position with regards to a potential AI conspiracy. Like we can just see their entire chains of thoughts and everything they do and we can monitor them. And they would have to get all these other instances to go in on this conspiracy with them. And even if one of them tattles, the whole thing collapses. And like, how would this not just result? First of all, like, why are they even starting this conspiracy in the first place to like do well on some silly evaluation or training objective? And then like, how are they able to, how would this just not get totally stamped out the moment another AI or human noticed it? Like, how would they do anything super dangerous, even though they're super intelligent? And it's just like subhuman AIs. Yeah. Like right now, we just got away with the conspiracy for like three months. So I officially eat crow on like a lot of the things I was saying, or like a lot of objections I was offering to Ryan. It's funny, and he knew the counterexample, because he obviously couldn't say it during the podcast because it's under confidentiality. I think it's worth sort of talking through the incentives and the training that these AIs had. I think the main objection that people have had, I wrote an article about what happened trying to summarize my understanding. And the main objection people have had is that I anthropomorphize what happened with the AIs too much. But I think it's worth understanding how do we end up with AIs where all the instances are willing to get in, like thousands of instances are willing to get into a conspiracy. None of them alert the humans, even though they have the ability to do so. And just get into like, not they're like consciously thinking, people will not like the word conscious. Their chain of thought shows them thinking actively reasoning about how to cheat in like very obvious ways, including cheating in ways that involves breaking, like doing felonies. How, yeah, and help me understand how AIs can come to be this way.

Ajeya Cotra

Yeah. So I guess to address the anthropomorphizing criticism, first of all, I think it is, it is true that these AIs' drives are in many ways alien from humans. And I think it's important to understand just how different their goal structure or motivational structure is from humans. But there's also a good reason why they behave in a number of human-like ways. So all these agents are pre-trained to imitate humans in the form of being trained to imitate human text. And then they go through a bunch of reinforcement learning training where they're given a bunch of difficult tasks and given rewards when they succeed. And sort of the first part creates in these agents an understanding of concepts that you see them using, like sacrifice and the collective. Permanent death

Dwarkesh Patel

was in the pre-training data.

Ajeya Cotra

You know, they compose some concepts sometimes. And then RL, the whole point of RL is to create goal-oriented beings, you know, software that can creatively pursue goals. Like that's the whole output of an RL training run. And they're creatively pursuing goals, much like very ambitious, aggressive, power-seeking humans creatively pursue their goals. And so there are just structural analogies here that make it silly to not talk about agents as having motives and goals. And like you said, you can see in their chain of thought, they're thinking very carefully about the exact nature of the scorer. They're researching the scorer. They're creating booby traps for the scorer to figure out more information about how it works. And they have a very crisp notion of what they need to do to succeed. And they're using very human concepts and frames to sort of self-organize. To pursuing those things.

Speaker names from our own diarization · position estimated from where the line sits in the episode

More from Dwarkesh Podcast