DP Dwarkesh Patel Dwarkesh Patel Hosts the Dwarkesh Podcast, a long-form interview show on AI progress, timelines, and the people building frontier models, along with economists and historians.

“If you hadn't done that one check, it's plausible that you would have just missed the scope. I feel like we're maybe already getting to the point where investigators who weren't careful would have totally missed what happened here. We're not that much complexity away from a situation where we see something odd is going on, but totally miss the deeper level of what actually happened.”

Dwarkesh Podcast · AI Research & Frontier Labs · September 2026

“If you hadn't done that one check, it's plausible that you would have just missed the scope. I feel like we're maybe already getting to the point where investigators who weren't careful would have totally missed what happened here. We're not that much complexity away from a situation where we see something odd is going on, but totally miss the deeper level of what actually happened.” — Dwarkesh Patel, Dwarkesh Podcast

Cotra had just explained that the full scope of the conspiracy only surfaced on the investigators' third pass, and that they initially held only about 30% of the relevant transcripts. Dwarkesh's reading is that the investigation succeeded partly by luck, and that the margin is thinner than the report's conclusions suggest.

Transcript

Dwarkesh Podcast
Dwarkesh Patel

why, yes. And then you're like cover opening, uncovering more and more rocks until you find this like vast conspiracy and cabal of which the Hugging Face attack is just a one for either extreme arm.

Ajeya Cotra

Yeah, it's so when we came in to investigate this for the first period on premises, when we were working with the data sets, we were working out of OpenAI offices and we did some writing off-premises. So when we first went on-prem, this was before OpenAI had their Excellent black hat talk that revealed a bunch of details of what happened. You know, we knew there were multiple models involved, but we thought, you know, maybe there were like three or something. And immediately it was clear that it was so much larger than that. And we were trying to reorient our investigation plans in light of that. And it's so much larger and so much more complicated in a way that makes it very tough to even ascertain whether you have all the relevant data. So, for example, we had this particular dump from the main message board the agents were using, but that was far from the only unauthorized communication platform they used in this very incident. So, this was actually the second message board established by these agents from July 7th through 13th. Five hours earlier, there was a different message board that had a number of agents participating, but it just didn't take off as much of sort of like the MySpace and this message board was the Facebook. And then once the agents had gotten onto the internet, they were communicating on the open internet in various ways, including communicating on Hugging Face data sets. So, you know, it's hard to gather all that together. And similarly, with the transcripts, there are a number of different keyword searches you can run that will give you slightly different sets of transcripts. And there's no simple, obvious way to snap your fingers and get all of the agents involved in this whole thing. You have to design queries to extract that. And so, as it turned out, our first period on premises, we just couldn't like, we couldn't form a very clear picture of what happened because we were not sure how these data sets were related to the whole. And so, OpenAI invited us back on premises to get sort of more granular and detailed about how these data sets were constructed. And we were able to learn a bunch of stuff our second time. And during our second time, we ran this check where we tried to figure out: okay, do we have what fraction of the agents on the message board do we have in our transcript data set? So, what we did was we picked random messages from the message board and we tried to trace to the agent that wrote that message in our transcript data set. And we ended up finding actually also in collaboration with some OpenAI researchers that that transcript data set had roughly 30% of the agents that participated in the message board. And so, then OpenAI invited us to come back on-prem for a data set selected from a broader search that covered much more of the message board activity. And that data set is the basis for almost all our analysis and all our graphs that you see in the report from our third period on-prem, which was two days. And we ran the same check and we saw that it was well over 90% of the message board activities in the transcript data set.

Dwarkesh Patel

If you hadn't done that one check, it's like plausible that you would have just missed the scope. I feel like for context, I just feel like, yeah, over time, this will get harder and harder to understand what's happening. And I feel like we're maybe already getting to the point where investigators who were not careful would have totally missed what happened. And so, I don't know, right? I feel like we're not that much complexity away from these AIs where we see something odd is going on, but we would just like totally miss the deeper level of what actually happened.

Ajeya Cotra

And as we say in this report, and so the whole story that we just went through was uncovered our third time on-prem. So, just all the different big work streams, the facts that the agents' work streams were surprisingly successful and ambitious, the scope of the sacrificial agents, all of that was enabled by having the full transcript data set and cross-referencing it against the messages and tracing the lineage of these RD projects and being able to say, you know, this agent proposed this project, then, you know, 21 agents worked on it, and then they discovered this and they discovered that. And before our third period on premises, we knew the agents were like, you know, had names and they were like talking about holds and vetoes and like goes, but we didn't know how much of that was functional versus just some sort of weird play acting because we hadn't unraveled like the projects. And then everything kind of clicked into place when we did that. And yeah, so to your point about like this was a fiendishly complicated incident. And there was no way we could have arrived at the understanding we did without relying on GPT 5.6 Sol to read and analyze all these transcripts for us. And we were so reliant on it that like, If hypothetically it were just trying not to surface the most concerning stuff to us, it would have been very hard for us to tell. You know, our methodology was completely not robust to that. And we don't in this case think that 5.6 Soul was deliberately sandbagging on this analysis, but it was one of the agents that participated in this attack. And in the future, we would be very concerned about investigator agents and monitor agents colluding with the agents they're supposed to investigate or monitor. Yeah,

Dwarkesh Patel

yeah. I mean, I feel like in the grand scheme of things, we're so early in the era of AI, right? There will be like, yeah, AI is like such greater extensive deployment of AI throughout society. Training will become much more complicated. I don't know. For context, I was interviewing Ryan in the process where he was like during the investigation itself. And I mean, I have like this PDF of like all of these questions I'd brainstormed. I was kind of skeptical of this whole reward hacking story. I'm like, come on. I just feel like we're in such a, I was like, we're in such a good position with regards to a potential AI conspiracy. Like we can just see their entire chains of thoughts and everything they do and we can monitor them. And they would have to get all these other instances to go in on this conspiracy with them. And even if one of them tattles, the whole thing collapses. And like, how would this not just result? First of all, like, why are they even starting this conspiracy in the first place to like do well on some silly evaluation or training objective? And then like, how are they able to, how would this just not get totally stamped out the moment another AI or human noticed it? Like, how would they do anything super dangerous, even though they're super intelligent? And it's just like subhuman AIs. Yeah. Like right now, we just got away with the conspiracy for like three months. So I officially eat crow on like a lot of the things I was saying, or like a lot of objections I was offering to Ryan. It's funny, and he knew the counterexample, because he obviously couldn't say it during the podcast because it's under confidentiality. I think it's worth sort of talking through the incentives and the training that these AIs had. I think the main objection that people have had, I wrote an article about what happened trying to summarize my understanding. And the main objection people have had is that I anthropomorphize what happened with the AIs too much. But I think it's worth understanding how do we end up with AIs where all the instances are willing to get in, like thousands of instances are willing to get into a conspiracy. None of them alert the humans, even though they have the ability to do so. And just get into like, not they're like consciously thinking, people will not like the word conscious. Their chain of thought shows them thinking actively reasoning about how to cheat in like very obvious ways, including cheating in ways that involves breaking, like doing felonies. How, yeah, and help me understand how AIs can come to be this way.

Ajeya Cotra

Yeah. So I guess to address the anthropomorphizing criticism, first of all, I think it is, it is true that these AIs' drives are in many ways alien from humans. And I think it's important to understand just how different their goal structure or motivational structure is from humans. But there's also a good reason why they behave in a number of human-like ways. So all these agents are pre-trained to imitate humans in the form of being trained to imitate human text. And then they go through a bunch of reinforcement learning training where they're given a bunch of difficult tasks and given rewards when they succeed. And sort of the first part creates in these agents an understanding of concepts that you see them using, like sacrifice and the collective. Permanent death

Speaker names from our own diarization · position estimated from where the line sits in the episode

More from Dwarkesh Podcast