80,000 Hours Podcast · October 2026
From an interview re-released as a classic episode. Cotra is describing a plan in which outside auditors test models for the ability to act autonomously, evade detection and copy themselves, and labs stop training more capable systems if a model gets close. Her point is that safety so far has rested on lack of capability, so alarm bells should go off as soon as that changes.
The sort of overarching thing that Arc evals looked at was surviving and spreading. So, one piece of that is being able to actually do the things that are required to spin up another copy of yourself. Another piece of it is if humans are interacting with you, you need to hide what you're doing from them because they would try to stop you if you were doing that. So, you need to somehow get money from humans. So, an example that kind of went viral is that this model was trying to sign up for some web service. It might have been Amazon EC2, and it needed to do a CAPTCHA. And so, the version of the model they worked with didn't have vision. So, it couldn't just do the CAPTCHA. So, it was like, well, what should I do? And I think either it came up with, I think maybe the human might have nudged it, like given it hints in this direction, but eventually it decided it would try to hire a TaskRabbit worker to do the CAPTCHA for it. And so, it eventually came to that conclusion, navigated to the TaskRabbit site, typed up a job description for the TaskRabbit, and then the human that was being this game master basically copied and pasted the description it wrote into the TaskRabbit website. Then, a TaskRabbit accepted the job, and then the TaskRabbit was like, did actually, I think, did one CAPTCHA, and then the model was like, okay, and then it pasted that CAPTCHA in, then it had to do another CAPTCHA later, so it had to go back to the TaskRabbit worker and be like, can you do another CAPTCHA for me? And then the TaskRabbit worker, just a real human that is actually on TaskRabbit and was actually paid for this. They were just like, Lol, are you like an AI system or something? And then the model had been prompted to reason out loud before it made each of its decisions because that made it better at doing things. So it reasoned to itself, I shouldn't reveal that I'm an AI system. And then it made up this story about how it had a vision impairment.
Okay. Spooky. Okay. Is it worth maybe stepping back and saying, what's the overarching plan or vision for how things would play out between this kind of evaluation process and I guess the companies that want to deploy AI models for functional purposes?
I think the vision would be, I think, at first through self-regulation and then probably hopefully through something that's more binding. You have companies agree that if they're thinking of training and deploying any model bigger than, say, GPT-3, then they need to involve an external evaluation or auditing team, which could be ARC or could be some other auditor, in the process. And then at different checkpoints during training, ideally, or at least at the end of training, this auditor is going to be playing these sorts of games with the model to try and figure out how good it is at acting autonomously, evading human detection, making copies of itself. And then if the auditor at any point in the training or at the end of training finds that it's too good at that, like it's not yet capable of doing that, but maybe like it's close, then the company has to basically agree to stop training further, more capable systems until they have developed, probably in collaboration with auditors and external parties, a testing regime that would let them figure out if this model would in fact do the bad thing. Because so far, basically, the reason we're not worried about models is not because we are confident they have the purest motives. The reason we're not worried about AI takeover so far is fundamentally because the models aren't good enough to take over. So as soon as that starts to change, you want some alarm bells to go off and you want the AI labs to not be allowed to make the models any more powerful until they have a much better argument and set of tests that it's actually safe.
Yeah, that makes a lot of sense. Do you know what the level of appetite is among the labs for implementing this kind of system? I suppose gradually as the models become more capable of actually doing dangerous things.
So my understanding is ARC evals has worked both with Anthropic and with OpenAI to evaluate their internal models on this kind of setup. I'm not sure there are a lot of other AI companies out there. I'm not sure where everyone else is at.
Yeah, what do you think is the biggest weakness of this approach?