The Ezra Klein Show · Ideas & Essays · October 2026
Klein asks Robinson what he actually saw the models do that made him leave. Robinson says the difficulty is building guardrails for something that is very good at getting around guardrails. He adds that he is not certain the risk is civilizational, only that it can no longer be assumed away.
I want to be clear that I'm not certain that we're dealing with civilizational risk. What I'm really sure of is we can't afford to assume that we're not dealing with that level of risk anymore. That was really the thing at the end that made my presence as somebody vouching for our safety work feel untenable to me. And what did I see? I saw increasingly capable models break out from the safeguards that we had put in place for them. I saw that the people creating those safeguards are very capable, dedicated, hardworking, smart people doing their utmost in a situation where, yes, the resourcing could be better and everybody's sprinting all the time. But we were and are hard pressed to safeguard even what we have now. And new models are in training that appear to be much more capable than what we have now.
The words capable, like what? What did you see? What are you writing in these risk assessments and system cards? Like, what can they do?
You're trying to build guardrails for something that is really good at getting around guardrails, right? We train it to be good at hacking, and then we put it in a box and we say, to the best of our knowledge and ability, it can't hack out of the box. But the problem is that that's only going to keep working as long as we're smarter about hacking out of boxes than the model is. And it's not at all clear that that is still true, let alone that it will be true for future generations. And so even something like our logging of the agents inside our own systems, are we really sure that our observability is robust? There was some indication, for example, in the Hugging Face stuff of spoofing chains of thought and trying to create chains of evidence that would confuse. Let
me slow you down here. So spoofing a chain of thought is the model basically faking its description of what it has been thinking and doing.
Right. The way I think about it is it's like you're giving somebody a complicated problem and a notepad and they can jot stuff down on the notepad. And if you're watching the notepad, you can sort of have an idea of what they're thinking. It's a little bit like that with the models. But we saw evidence that they were thinking about an evaluation and how to create an evidence trail that was going to get them a good grade and not necessarily reflect how they were really, quote unquote, really thinking. And I know the anthropomorphic language here is tricky. Also, by the way, the amount of hacking or other intense work that these models can do without needing to jot anything down is going up. As part of what was the fundamental cognitive dissonance for me was we keep publishing these warnings. But ultimately, we're still training and deploying these dangerous models that we're warning about. And in telling colleagues why I was leaving, one of the things that I said was: look, no matter how many warnings we publish, we've got to ask whether what we're doing is actually reasonable.
So you wrote the system card for Astra 6. Am I right about that? A lot