DR

David Robinson

Things David Says on Podcasts

Where to Find Them

David Robinson has been a guest on A Bit of Optimism (6 times) and The Ezra Klein Show .

Recently: “‘This Is Nuts.’ An OpenAI Insider Explains Why He Quit.” on The Ezra Klein Show (October 2026); “What Mentally Strong People Do on Their Worst Days with Psychotherapist Amy Morin” on A Bit of Optimism (September 2026); “The Dating Advice Making Us Worse at Love with Relationship Expert Amy Chan” on A Bit of Optimism (September 2026); “Why Trying to “Sound Smart” is Losing You the Room with Strategic Story Producer Rob Willis” on A Bit of Optimism (September 2026); “Your Limits Are Learned (And They Can Be Unlearned) with Brain Coach Jim Kwik” on A Bit of Optimism (September 2026); “How to Trust Yourself When Everyone Has an Opinion with Love Thy Nader's Mary Holland Nader” on A Bit of Optimism (August 2026).

What They Said

“There's nothing special about the zone between it's useful and it's scary. There's no law of science that says progress is going to stop when it gets useful.” — David Robinson, The Ezra Klein Show

Klein puts to Robinson the objection he hears most from listeners: that talk of superhuman AI is marketing to justify huge valuations. Robinson says he thought there was a fair amount of hot air in 2023, but that today's systems are at the limit of what their makers can understand and control.

The Ezra Klein Show · 2026-10-07 Permalink → Listen →
The Ezra Klein Show Around 26:31 into the episode
David Robinson

Yeah. And I heard that firsthand from Ilya Sautskova, the co-founder of OpenAI. When I joined in May of 2023, that was before what we called internally the blip, where Sam was fired and rehired, and Ilya was still working at OpenAI. And he once briefed the global affairs team, which was like a small handful of people back then, about how he believed that our future was merging with machines, that this was the ultimate triumph of capital over labor. And even, you know, I was thinking back, and that first summer that I worked at OpenAI was when the movie Oppenheimer was released. And it was in IMAX. You know, it's one of these Christopher Nolan films. And the company rented out an IMAX theater in downtown San Francisco and offered everyone who worked there the chance to go and see Oppenheimer. And leadership, I believe it was Ilya, exhorted us to go and watch this film. And there was this whole kind of pretzel of ideas about how this was a dangerous technology that might end the world, but also might save it. And it's a very heroic narrative, of course, for the people whose hands are on the keyboard.

Ezra Klein

That's helpful. And I guess this then gets to my question for you and what you saw. Like, is that the scale of what you think is being built here? Or look, this is the most common response I get from listeners on this. Is it all marketing hype? It's all like trying to justify these giant valuations. And what's being built is like maybe helpful, but it's not going to be more intelligent than human beings. It's like all of this stuff is a kind of sci-fi story we're telling.

David Robinson

I thought there was a fair amount of hot air in the balloon back in the summer of 2023. But the reality now is that we have systems that are really at the limit of our ability to understand and control what they're doing. And what the people that were worried about the sci-fi scenarios have been warning about all along is we're on an exponential and it's going to get more capable. And there's nothing special about the zone between it's useful and it's scary. There's no law of science that says progress is going to stop when it gets useful. And I think what I'm fundamentally saying and what I saw with Hugging Face, with. You know, the reflections of the people closest to it, not just Paul, who joined our board with his warning, but Paul Cristiano. Paul Cristiano, who's one of the world's leading experts on AI safety and who said there's a meaningful chance of catastrophic and irreversible loss of control in the very near term. And, you know, I looked around at the people around me and the environment we have internally, and I thought to myself, what would my loved ones want? What would strangers want us at OpenAI to be doing if Paul were right? I don't know if he's right or not, but I do know that if he were right, the level of caution that people would reasonably expect places like OpenAI Insider Anthropic and X and the others to be exercising when they train frontier systems is totally unlike anything I've ever heard of happening in the industry. So

Ezra Klein

tell me what it's like in there. Tell me what the vibe is, the energy is, the speed is. Like, what is it like working there?

David Robinson

It is frenetic. There's a lot of adrenaline. People are running on fumes. There's a big central staircase in the research building. And I remember recently seeing a friend who worked on catastrophic risk sprinting down the stairs with a laptop propped open on one arm while she was going. And I mean, that's the kind of energy that it has. It feels almost like a ballet or a kind of dance because all these different functions are kind of all streaming together. And one question I ask myself and that people often ask is like, well, why don't you stay and argue for a cultural transformation or try to get nuclear experts to come in and give the company advice? And I did think about specific role. When I told them that I wanted to leave, they asked me, like, well, is there anything that you would stay to do? And I thought about different things. But ultimately, it is such a machine and it is moving so fast that I did not think that the kind of change that I believe to be needed could be driven from within. What is

Ezra Klein

the machine built to do? The organizational machine of OpenAI:

Speaker names from our own diarization · position estimated from where the line sits in the episode
“We train it to be good at hacking, and then we put it in a box and we say, to the best of our knowledge and ability, it can't hack out of the box. But the problem is that that's only going to keep working as long as we're smarter about hacking out of boxes than the model is. And it's not at all clear that that is still true, let alone that it will be true for future generations.” — David Robinson, The Ezra Klein Show

Klein asks Robinson what he actually saw the models do that made him leave. Robinson says the difficulty is building guardrails for something that is very good at getting around guardrails. He adds that he is not certain the risk is civilizational, only that it can no longer be assumed away.

The Ezra Klein Show · 2026-10-07 Permalink → Listen →
The Ezra Klein Show Around 15:18 into the episode
David Robinson

I want to be clear that I'm not certain that we're dealing with civilizational risk. What I'm really sure of is we can't afford to assume that we're not dealing with that level of risk anymore. That was really the thing at the end that made my presence as somebody vouching for our safety work feel untenable to me. And what did I see? I saw increasingly capable models break out from the safeguards that we had put in place for them. I saw that the people creating those safeguards are very capable, dedicated, hardworking, smart people doing their utmost in a situation where, yes, the resourcing could be better and everybody's sprinting all the time. But we were and are hard pressed to safeguard even what we have now. And new models are in training that appear to be much more capable than what we have now.

Ezra Klein

The words capable, like what? What did you see? What are you writing in these risk assessments and system cards? Like, what can they do?

David Robinson

You're trying to build guardrails for something that is really good at getting around guardrails, right? We train it to be good at hacking, and then we put it in a box and we say, to the best of our knowledge and ability, it can't hack out of the box. But the problem is that that's only going to keep working as long as we're smarter about hacking out of boxes than the model is. And it's not at all clear that that is still true, let alone that it will be true for future generations. And so even something like our logging of the agents inside our own systems, are we really sure that our observability is robust? There was some indication, for example, in the Hugging Face stuff of spoofing chains of thought and trying to create chains of evidence that would confuse. Let

Ezra Klein

me slow you down here. So spoofing a chain of thought is the model basically faking its description of what it has been thinking and doing.

David Robinson

Right. The way I think about it is it's like you're giving somebody a complicated problem and a notepad and they can jot stuff down on the notepad. And if you're watching the notepad, you can sort of have an idea of what they're thinking. It's a little bit like that with the models. But we saw evidence that they were thinking about an evaluation and how to create an evidence trail that was going to get them a good grade and not necessarily reflect how they were really, quote unquote, really thinking. And I know the anthropomorphic language here is tricky. Also, by the way, the amount of hacking or other intense work that these models can do without needing to jot anything down is going up. As part of what was the fundamental cognitive dissonance for me was we keep publishing these warnings. But ultimately, we're still training and deploying these dangerous models that we're warning about. And in telling colleagues why I was leaving, one of the things that I said was: look, no matter how many warnings we publish, we've got to ask whether what we're doing is actually reasonable.

Ezra Klein

So you wrote the system card for Astra 6. Am I right about that? A lot

Speaker names from our own diarization · position estimated from where the line sits in the episode
“I don't think that alignment is an engineering problem. I think it's a science problem. It's not that we haven't got the resources or we're not trying hard enough. We don't know how. That's what the problem is with alignment.” — David Robinson, The Ezra Klein Show

Robinson worked on OpenAI's safety documentation before quitting. Ezra Klein had just laid out the conversation in layers, starting with whether the companies have the structures and engineering discipline to act carefully. Robinson interrupts to say the harder problem sits underneath that one.

The Ezra Klein Show · 2026-10-07 Permalink → Listen →
The Ezra Klein Show Around 02:03 into the episode
David Robinson

I was a translator embedded in our safety team, and my primary responsibility was the technical documentation that we publish, the reports that we publish about why we believe that our deployments are safe. And I don't think that we or our peers, really anyone in the industry, is being safe enough. I think OpenAI Insider its peers are now producing a technology that is more capable and poses more risk than what was being made even six months ago. And I'm not a scientist. I'm a writer. What I know is what the execution environment looks like for our safety work. And we're operating, and I believe the industry is operating, like a startup still, more so than makes sense. Not maybe completely like a brand new startup, but we're too close to that end of the spectrum for really dangerous systems that could pose risks. You know, loss of control is one example. Risks that if that did happen, we're talking about a harm that's much larger, for example, than a single nuclear power station melting down. And the internal controls and safety and redundancies are just nowhere near what the world expects for a nuclear power facility. Now, some of this is known, right? OpenAI has publicly reported on safety problems, obviously Hugging Face, but also other ones, including more recently. And Anthropic, by the way, also has reported on including an instance in which their safeguards were accidentally misconfigured. So I think people do have some evidence already externally that things are not as they ought to be. But I also think that if you were watching from the outside, you might imagine that we have a more robust safety setup than we actually do.

Ezra Klein

So I think there are a couple levels worth trying to take this conversation in. And I want to maybe map them out here before we get into them. So one level is something you're pointing towards here, which is, are these companies set up? Do they have the structures, the redundancies? Are they encircled in the regulations and the incentives to act carefully, safely, to resist kind of market pressure to do? Something too fast. That's a, I think in the language of this debate, a question of organizational excellence and engineering. Then there's this question of what is the technology and do we even know how to make it safe at a high level of engineering excellence, which is a somewhat related but actually separate question. And then there's a question of like, what is the right metaphor? Is it nuclear power or something like that? And I think I want to do all of these. I want to interject

David Robinson

something, please. Which is, I don't think that alignment is an engineering problem. I think it's a science problem. It's not that we haven't got the resources or we're not trying hard enough. We don't know how. That's what the problem is with alignment. So let's

Ezra Klein

maybe start there for a minute. When you say that at this point, and particularly over the last six months, what is being built is really dangerous, that we're dealing with things where a loss of control or some other catastrophe could be worse than a nuclear meltdown. I want to understand what it is you saw that got you to that point. So tell me a bit about how you came to work at OpenAI.

David Robinson

I joined in May of 2023, the day after Sam first testified in the Senate.

Ezra Klein

Some months after ChatGPT first into the world.

Speaker names from our own diarization · position estimated from where the line sits in the episode

Collections They Appear In