← Things People Say on Podcasts

“this might be the clearest warning shot we ever get for loss of control because these agents…did a bunch of extremely sophisticated, difficult things to pursue this relatively long horizon cheating goal, but they seemed completely uninterested in covering their tracks from humans.” Ajeya Cotra · Dwarkesh Podcast

AI Research & Frontier Labs · September 2026

“this might be the clearest warning shot we ever get for loss of control because these agents…did a bunch of extremely sophisticated, difficult things to pursue this relatively long horizon cheating goal, but they seemed completely uninterested in covering their tracks from humans.” — Ajeya Cotra, Dwarkesh Podcast

Cotra co-authored the investigation into the agent swarm that compromised Hugging Face, and this is her closing thought on the episode. Her argument is that the combination was unusually legible: capable enough to pursue a long-horizon goal, careless enough to leave the evidence in plain sight. She expects future agents to be far more attuned to the fact that humans are watching.

Transcript

Dwarkesh Podcast
Ajeya Cotra

time now to develop remediations than we will then. So the sooner the better, I think, to understand what's going on.

Speaker 3

Yeah. I think that might be a good place to close. I don't know if there's anything more that you feel like might be worth saying. I

Ajeya Cotra

think one thing that feels especially concerning to me about this whole incident is that this might be the clearest warning shot we ever get for loss of control because these agents were in like this interesting middle ground of they did a bunch of extremely sophisticated, difficult things to pursue this relatively long horizon cheating goal, but they seemed completely uninterested in covering their tracks from humans. And as we discussed, it seems possible that even this incident would have been like much less striking in some ways if these AI systems had had it made salient to them that they should sort of act more quietly because humans might watch them. They might have gone much further than they did, especially after the 19th. But then future AI agents, for a number of reasons, will probably be much more attuned to the human world. Maybe companies will start instituting online training that includes sort of aspects of human judgment in the evaluation. Or maybe agents just generalize further and just like more keep it in mind that humans are out there in the world and they're sort of potentially a resource or an obstacle. And so even if future AI agents do much more deeply damaging things, like institute a persistent covert rogue deployment inside a company and siphon off its compute resources and poison the training data of future models, it is not at all obvious that that will penetrate out into the outside world in this dramatic way, resulting in the hack of some external party that wasn't actually, frankly, very important for these agents to hack and then lead to this moment where there's some investigation. And even if investigation is done, it might be much harder to make sense of what happened if agents were systematically compromising telemetry and logging like these agents tried to do, but didn't succeed at. And if these agents stop thinking in English and start thinking in neural activations, then you'd be forced to just ask another AI agent what was happening and have no way to cross-check it against the ground truth. And so, you know, I think that probably much more concerning things will happen, but it may never be as clear as this before it's far too late. Yeah.

Speaker 3

And again, you're too humble to say this, but like it is very likely would not have been clear even in this case, were it not for you and Yalmar and Ryan sort of getting maybe that third shot on goal with the investigation you did and then both your technical competence and things lining up in the right way such that you could have like done like even this investigation, which in retrospect we could understand and is like a parsable scheme or whatever, is sort of contingent that we actually got this warning shot. And it's just not that complicated in the grand scheme of things. Again, we're so early in the story of AI.

Ajeya Cotra

And this is such a small number of agents compared to everything that's going on across all the frontier AI companies right now, let alone

Speaker 3

a year

Speaker names from our own diarization · position estimated from where the line sits in the episode

More from Dwarkesh Podcast