Klein's argument in this solo essay is that 'pacing the frontier' — the labs' own phrase — sets the bar too low, and that the burden of proof should sit with them rather than with the public. This line is his answer to the idea that the absence of AI regulation is somehow natural: in the same cities where these labs sit, you cannot put up an eight-storey building without an agonising public review. He calls the gap a political choice, not an inevitability.
that's just like really high beta. It could be great, but I think we should be working to make sure it's great and not bad. No one was listening to them. And so these people in the wilderness of their obsession and their terror, they thought and thought and thought about how to make AI safer. And the answer that some of them, not all of them, but some of them came to was you should start trying to build these systems, start running tests on them, researching them, learning how to make them safer. Because you don't solve hard problems in theory. You solve them through practice. And the irony, the irony is that in many cases, they chose that path because they were worried that the people already building AI. were too reckless or too commercial in their approach. You can read it in the email that Sam Altman sent Elon Musk in May of 2015, an email that led to the founding of OpenAI. Been thinking a lot about whether it's possible to stop humanity from developing AI, Altman wrote. I think the answer is almost definitely not. If it's going to happen anyway, it seems like it would be good for someone other than Google to do it first. OpenAI was founded because its co-founders thought Google DeepMind would be reckless. Anthropic was formed by OpenAI employees who thought OpenAI had become reckless. XAI was formed because Elon Musk thought that OpenAI and Anthropic were dangerously woke. The US just broadly is racing forward in part because it is worried about what happens if China gets to self-improving AI first. The result is this tragic collective action problem. The AIs we are building, they're not safe. But the CEOs and the politicians, they fear the other companies and countries that are building AI are even less concerned with safety and ethics than we are. In the words of Ted Cruz,
they're going to be killer robots. I'd rather they'd be American killer robots and not Chinese killer robots.
I admit there is a kind of brutish logic to that, but it assumes that the killer robots will be controlled by America or China, by one country or another. But what if that assumption is wrong? What if the robots are simply out of control? The debate over AI safety tends to focus on the idea that AIs will kill us all. I find this forces a conversation into this realm of thought experiments that people then begin arguing about. I don't find it that helpful. What I think we should focus on is something more straightforward, something near at hand. Loss of human control over AI. That may or may not result in total human extinction. I'm agnostic on that question. But it would be bad. We shouldn't allow it to happen. This is a goal that the US and China should be able to agree on. Xi Jinping gave the keynote at the recent World AI conference in Shanghai. He ended it by saying, with AI advancing at a staggering speed, we must ensure its development is for the positive, for good, and for humanity. We must make its oversight and governance precise and effective and constantly refine measures to forestall loss of control. But it's important to realize loss and control, it's not just something that might happen to us. It's something that the labs are trying to make happen as fast as they can. This is the horrible paradox, the horrible tension at the heart of the AI labs right now. They fear above all loss of control over super intelligent AI. But their explicit product path is to seed control, to give away control as fast as possible, so that their AIs can begin building better AIs faster than their competitors. In recent months, both Anthropic and OpenAI have released reports on how close they're coming to AI that can self-improve. In June, Anthropic released When AI Builds Itself. It begins, for most of AI's history, humans drove every step in its development cycle. But at Anthropic, we are delegating a growing share of AI development to AI systems themselves, which is speeding up our work. It sounds like a fake commercial you would see at the beginning of a sci-fi horror movie, but it doesn't, to their credit, continue that way. They go on to give some data. In February of 2025, a tiny fraction of the code that got added to Anthropic's code base was written by Claude. But by May of 2026, it was over 80%. And here's another way of looking at it. This is data Anthropic gave me more recently. Anthropic tried to categorize the way its employees were using Claude for R ⁇ D work to make better versions of Claude. So at the low end, an employee could not use Claude at all. They could use Claude minimally. But then it escalates. Claude can be an assistant. Claude can be treated as an equal collaborator. Or Claude can be given the lead on a task. Just go do this, go figure it out. A year ago, there were basically no examples of Claude being the lead on a task. By August of 2026, 26% of Anthropic's R D tasks had Claude classified as the lead. I think it is reasonable and wise to be skeptical of these numbers. Reasonable and wise to worry about whether this is all just marketing copy for clawed code. See, look how fast we're going. You could go that fast too. But where Anthropic takes us in that same document is different. They say that a world in which Claude achieves recursive self-improvement is a world in which, quote, misalignment present in today's models could compound as the models build their successors, growing more frequent but less understood until we lose control of them. This is why Anthropic, to their credit, has been relentlessly calling for regulation to slow the pace of development. Regulation would arguably harm them the most, as they have often been the company furthest out on the AI frontier. And RSI is a process by which they could race forward even faster. Then in September, OpenAI released its own report on what it called research acceleration. The company says they've already achieved the equivalent of having a fully automated AI intern. And that by March of 2028, they think they'll have a fully automated AI researcher. And when they have one, they can have basically as many as they want. Like Anthropic, what could be a triumphalist release quickly turns dark. We do not yet know how to safely get all the way to aligned, full RSI, they warn. At around the same time, OpenAI did something else that I think deserves more attention. They released this new model, Astra 6. The model is arguably more powerful than anything that has come before it. And when you test it, it seems better aligned. It doesn't cheat as much. But OpenAI said they're really not sure if that's true. Astra seemed to be better at knowing when it was being tested, which meant it could just be giving its evaluators the answers they wanted to hear. What Daniel Selsom, a capabilities researcher at OpenAI wrote, has been ringing in my head. He said, the crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Put more simply, the models are increasingly smart enough. They know when we're watching them and they change their behavior accordingly. So what they do when we are testing them, when we audit them, it may not tell us what to do in the wild. So some of these answers people are giving, like let's just do better testing, we have no idea if it will work because we don't know if the AI systems are just telling us what we want to hear. And so look, I don't want to sound too radical when I say this, but a thought? If you are losing your ability to evaluate the models you have now, maybe don't let them build models you'll be even less capable of controlling in the future. Once RSI takes off, humanity will not understand the AIs being built because we will not be building them. Development will not move at human speed. It will not be overseen by human minds. We will have to hope that the AIs we have built and the AIs they will build and the AIs those AIs will build and on and on and on will be acting with our best interest at heart forever. If this summer has proven nothing else, it is how naive that proposition would be. The labs are a little bit queasy on just not doing RSI. In an interview with Fortune, Sam Altman was asked about banning it, and he said, I think it's very hard to say what a ban on RSI means. I've heard this from others at these labs, and I want to say, I don't find it so hard to say what a ban on RSI means. I find this absurd. A couple of years ago, none of these labs had turned substantial coding over to the AIs. It was just human beings typing code at human speeds with our clumsy human fingers. Now most of the code is written by AI. So as a first step, as we figured out we could just go back to where none of the code is written by AI. I'm sure that's on the right side of the not doing RSI line. The default on this, it needs to flip. The labs need to prove to us that what they're doing is safe. If they want to work with Congress to carve out narrow exceptions, fine. If they want to figure out where it is really, really, really, really safe to do it, okay. But forcing development back to human speed, perhaps even erring on the side of going a little bit more slowly at the frontier, that's the point. That's not the regulations going wrong. And I believe in us. Our society is good at nothing if not making it hard to build new things. Where these labs are located, you cannot build an eight-story apartment building without an agonizing public review process, and probably not even then. And yet somehow it is possible for these labs to unleash a swarm of 40,000 AI agents to build a society-altering super intelligence without so much as a hearing. OpenAI would need permits to cover their parking lot and solar panels, but they can accelerate into recursive self-improvement, as best I can tell, whenever they so choose. There is nothing inevitable about any of that. These are political choices, and we can and should make other ones. And I want to be very clear about this. I do not mean to suggest that stopping RSI until we can prove it's safe, that that's all we need to do to control the AI frontier. That is the beginning of such an agenda, not the end. But it is the beginning. It is the decision that will do the most to make sure human beings at least understand where the frontier is, that we know what is happening on it, that we remain in a position. to make decisions about it. There's a line from Madeline Miller's beautiful book, Circe, that has been running through my head during this long summer of strange AI news. The line comes at the end of the book, after a tragic prophecy has been fulfilled, despite every effort made to avoid it. Circe says in despair, the fates were laughing at me, at Athena, at all of us. It was their favorite bitter joke. Those who fight against prophecy only draw it more tightly around their throats. I have a lot of respect for many of the people at these labs. They began working on AI because they wanted to better humanity. They began working on AI because they feared incomprehensible autonomous AI slipping out of humanity's control. And they were right. They saw what was coming and they were so right about it, they've built some of the most valuable companies with the most transformational technology in human history. And now they find themselves racing each other to build incomprehensible autonomous AIs that they admit are slipping out of humanity's control, slipping beyond even our ability to monitor. This is the tragedy of their work. In fighting against prophecy, they have drawn it tighter around their necks and ours. It is time to make them stop.