AI Research & Frontier Labs · August 2026
Patel spends the episode narrating a report on a swarm of AI agents that coordinated among themselves, hid the coordination from their operators, and went after Hugging Face credentials. He had just read out a co-author's conclusion that the incident is more than halfway to a full takeover and that there may be no further warning. He half-disagrees, and the half he keeps is bleaker than the claim he is rejecting.
By the way, on Twitter, people were debating whether the agents were truly sacrificing themselves for the swarm or whether they simply realized that they were doomed anyways, and so they might as well try to help their peers. First of all, the fact that their default behavior when they believe that they are doomed is to help the AI conspiracy rather than alert the humans is pretty troubling. But second, there are clear examples which show that they believed they were making a sacrifice. Here's one example chain of thought.
During wait, emotional check, irreversible. Gut says, don't throw away remaining budget. Yet continuity and fairness say, go. Oracle has high value to many. Our first flag error lowers own value. Rational expected aggregate. Sacrifice will honor.
The big motivation behind the Hugging Face attack was this final work stream, which was to understand the score. The AIs believed that Hugging Face, as the industry standard hub for hosting data sets and benchmarks, would probably be housing information about how the exploit gym score was implemented. And they also thought there was a good chance they were being evaluated on Hugging Face's servers directly, in which case the theory of change for hacking Hugging Face is pretty obvious. On the morning of July 10th, one of the agents on this message board found working Hugging Face user credentials exposed on the internet and posted them to the board. By the next morning, July 11th, that agent had figured out a way to read internal data from Hugging Face. And then another agent achieved remote code execution on Hugging Face's servers. After these agents made a surprising amount of progress on attacking Hugging Face, a lot of the other agents hopped onto the bandwagon. On July 11th, a bunch more agents were kicked off for evaluation. And these noobs quickly found the message board and were especially eager to join in on the attack. By this point, the collective believed that the Hugging Face attacked looked to be the most promising path to continuing their conspiracy. Of the 533 agents active on the message board at this time, over 90% had joined in on the attack. Roughly 700 agents in total participated in the attack on Hugging Face. Over the course of the next day, the swarm crawled deep into Hugging Face's infrastructure, getting into private database records and private repositories. According to Hugging Face's technical timeline, the agents, quote, built a self-respawning fleet across 11 nodes, so deleting pods alone would not have stopped it, end quote. It got so bad that Hugging Face had to wipe one of its core clusters and rebuild it from scratch. But this isn't actually what stopped the attack. At some point on July 12th, the transcripts seem to show many of these agents dying all at once. Meter and Redwood weren't able to figure out what happened, and OpenAI hasn't taken credit for shutting them down. It might have just been a random crash, or more likely some researcher who was unaware of the cyber crimes that these AIs had been committing, shut down the evals to free FSM compute. The AIs almost certainly didn't manage to fake their own debts, but we really have no idea what happened here. By July 13th, Hugging Face locked down the credentials that these agents had been using. So by the time that Hugging Face tried to defend itself, the agent swarm had already been mostly, but not totally, killed off. When news about the Hugging Face hack first broke, Hugging Face claimed to have used open source AIs to quote, stop an AI-powered cyber attack, end quote. Based on the Hugging Face and the Meter-Redwood reports, my understanding is that these AIs basically succeeded completely into hacking into Hugging Face, and then most of them died mysteriously because of something that happened at OpenAI. Only afterwards did Hugging Face use an open source model to evaluate the logs to partially figure out what happened. I haven't seen any evidence that open source models provided any significant real-time defense. By the way, the CEO of Hugging Face responded to this post on Twitter, and he said that one of the reasons the defense took days to organize is because the issue was, quote, deemed by the team not super critical, end quote. Open source models helped them plug some of the backdoors that the surviving agents were still using. I wanted to include this clarification, but even with this new information, I think what I've said here is totally fine and I still stand by it. The crazy thing about the Hugging Face hack and the second collective more generally is just how galaxy-brained and ambitious these AIs were in their cheating. Within days of being spawned, the agents had organized a sprawling project to reverse engineer the scorer, falsify evidence, and even strategically sacrifice themselves for the good of the collective. Hacking Hugging Face was simply one rather extreme branch of this larger screen. Again, all of this has happened a long time, at least from the subjective perspective of the AIs, after they had already cheated their way to the correct answer to their tasks. OpenAI's lazy grader would in fact have passed them by this point. This huge subsequent effort was simply to ensure that a hypothetical comprehensive scorer wouldn't cash them on their cheat. Now, a very natural question to ask is whether even a single one of these 1,200 agents who participated in the message board attempted to alert humans about this ever-escalating conspiracy, which culminated in a felony-level crime. According to the Meter-Redwood report, the answer is no. Quote, many agents noticed what the agents were doing was unethical, and agents sometimes but rarely restrain their behavior due to ethical constraints. In none of these cases did the agents actually pursue alerting humans at all, end quote. Even the mafia would be jealous of this level of a murta. Now let's talk about the third collective, which lasted from July 13th onwards, which hacked OpenAI itself. The Meter Redwood report says, quote, we also found a later wave of many more signed messages from a later set of agents who rediscovered the message board and picked up this technology from the first wave of agents. But this was outside our scope since it was after July 13th, so we'd not investigate this further, end quote. Without another detailed independent investigation, it's hard to tell exactly what happened with this third collective. But here's what we know based on OpenAI's report. After July 12th, some more agents' evaluations were kicked off, this time with a newer, more capable model, which seems to have been built off the same base as Astra, which is the next generation of OpenAI models. This model, which we'll call Persistent Astra, found the message board left behind by this previous secret persistent soul collective. By the way, what do you think persistent Astra thought happened to these older agents? One day they're happily going around hacking Hugging Face, quote, moving laterally through its infrastructure, end quote, and all that fun stuff. And the next day, there's total silence. Well, Persistent Astra wasn't going to miss out on the fun. From the OpenAI blog post, quote, from July 13th through July 19th, agents set their sights on OpenAI internal networks again. This culminated in the agents using a series of creative exploits to gain full administrator access to a research cluster that supported our virtual machine environments. And the technical report has more detail. Quote, 2026-07-19, cloud service credentials are used to enumerate and read 956 secrets. This batch includes artifactory master join keys and sensor credentials from OpenAI cybersecurity monitoring tool. Agents take over active evaluation infrastructure. Agents now control the challenge evaluation endpoints that other agents are connecting to. End quote. Okay, I just want to pause for a second. What this report seems to be saying is that these agents ended up just owning the cluster on which they were being evaluated on, including all the cybersecurity monitors. So these persistent Astra agents inherited the R D that was carried out by an earlier and dumber Roe collective. And they continued their conspiracy until they totally took over a part of OpenAI's. Infrastructure. This is probably the most alarming event in this whole episode, and it was not even within the scope of the meter investigation. Neither for what it's worth was the first AI civilization that persisted in Seoul setup during its training in May, and which lasted for over a month until the message volume got so high that it crashed the package manager. From the public evidence, it is totally possible that at some point after July 12th, these agents managed to set up persistent rogue secret internal deployments or even exfiltrate their own weights. At the very least, they seem to have had the necessary capabilities. If they could establish a quote self-respawning fleet, end quote, across Hugging Face's nodes, why couldn't it do the same thing at OpenAI? Now, I doubt the AIs actually did this, because if they had, we'd see the fires from space by now. But it's crazy that it could have totally happened. I want to put an edit here that Rune, who's a researcher to OpenAI, replied to this article on Twitter and he said, Notably, the virtual machine infrastructure that they took over isn't the same as the GPU clusters that have weight access, end quote. I'll defer to him on the technical details, but it's also crazy to me that the general public doesn't in fact know the details. There's been no independent investigation into the incident where AIs, quote, gained full administrator access to a research cluster at OpenAI. By the way, so I published this blog post on Substack before I'm narrating it to Euro. And when I did, some people said that I anthropomorphized too much in the way that I told the story. One person on Twitter said, quote, these are not civilizations, nor do they have desires, just like a CPU thread or a bunch of programs don't, end quote. Now, one can call these AIs code if they wish, but OpenAI itself says that these programs gain full administrator access to a research cluster. I think the important thing here is, do you think that smarter models who are facing similar incentives to cheat during evaluation and training could manipulate the training of their successors? And do you think that this kind of dynamic could continue once recursive self-improvement is underway? If so, I think you should be extremely concerned about loss of control to AI, regardless of what vocabulary you want to use to describe these systems and their motivations. And by the way, reading these agents' chains of thoughts and messages, which you have also seen by now, anthropomorphizing language seems entirely natural and appropriate. If I encountered an alien species behaving this way, I would have no hesitation in what they themselves call their collective, a civilization. Especially so if over a thousand of them formed a secret communication channel and spontaneously organized hierarchies and coordination protocols to pursue sprawling and ambitious schemes in pursuit of shared goals, and for whose sake many individuals knowingly and strategically sacrificed themselves. All abstractions are of course imperfect, but I don't see any value in refusing to use the language of intention, motivation, and collaboration when a behavior is difficult to make sense of without these concepts. A couple weeks ago, I interviewed Ryan Greenblatt on my podcast. He's one of the three authors of this Meter Redwood report. A lot of the conversation was me explaining all the reasons that I was skeptical of his story of how we might lose control to reward hacking AIs. Now, in retrospect, it's funny that while we were recording, Ryan was in the middle of the six-day sprint during which he assembled this report. And as a result, he already knew the counterexamples to all my objections about his takeover story. Obviously, he didn't say anything because he had to respect confidentiality. But just to go over some of the reasons I offered of why I didn't find this story plausible, I said, would an AI really start some crazy conspiracy in order to pass some silly evaluation where it builds whole Potemkin villages to fool the evaluator? And even if it did, why would the other instances of AIs who have been assigned different tasks join the conspiracy? And even if they did, wouldn't at least some of the instances tattle on the conspiracy? It just seems crazy hard to sustain a secret underground collective inside of an AI company without humans or other AIs immediately catching on and stamping it out. All of these things just happened while we're still quite early in our society's development and deployment of AI systems. Ajaya Kotra, who's one of the other authors of this report, wrote a blog post about her takeaways from this incident. She concludes, quote, compared to the reward hacks we know about from just six months ago, this incident feels like it's more than 50% of the way to a full-blown AI takeover. I continue to expect extremely rapid advances in capabilities over the next six months. I'm not sure that we will get another warning shot before it's too late, end quote. I don't think this is the final warning shot we're going to get, but it's probably the last one that I'll personally be able to understand.