80,000 Hours Podcast · October 2026
From an interview re-released as a classic episode. Cotra calls this one of the biggest things she worries about. If catching an AI system doing something sneaky and penalising it reliably fixed the problem, developers could work iteratively and empirically. The trouble is that a model that has learned honesty and a model that has learned to hide better look the same from outside.
That's right.
How big a problem is this?
So I think it is one of the biggest things I worry about. If we were in a world where basically the AI systems could try sneaky, deceptive things that weren't totally catastrophic, didn't go as far as taking over the world in one shot. And then if we caught them and basically corrected that in the most straightforward way, which is to give that behavior a negative reward and try and find other cases where it did something similar and give that negative reward, and that just worked, then we would be in a much better place because it would mean we can kind of operate iteratively and empirically without having to think really hard about tricky corner cases. But if in fact what happens when you give this behavior a negative reward is that the model just becomes more patient and more careful, you'll observe the same thing, which is that you stop seeing that behavior. But it means a much scarier implication.
Yeah. It feels like there's something perverse about this argument because it seems like it can't be generally the case that giving negative reward to outcome X or kind of process X then causes it to become extremely good at doing X in a way that you couldn't pick up. Like most of the time when you're doing reinforcement learning, don't you kind of, as you give it positive and negative reinforcement, it tends to get closer to doing the thing that you want. Do we have some reason to think that this is an exceptional case that violates that rule?
Well, one thing to note is that you do see more of what you want in this world. So you'll see perhaps this model, instead of writing the code you want it to write, it went and grabbed the unit tests you were using to test it on and just like special cased those cases in its code because that was easier. It does that on Wednesday and it gets a positive reward for it. And then on Thursday, you notice, oh, the code totally doesn't work and it just copied and pasted the unit tests. So you go and give it negative reward instead of positive reward. Then it does stop doing that. Like on Friday, it'll probably just write the code like you asked and not bother doing the unit test thing. So this isn't a matter of reinforcement learning not working as normal. I'm kind of starting from the premise it is working as normal. So all this stuff that you're whacking is getting better. But then it's a question of what does it mean? Like how is it that it made a change that caused its behavior to be better in this case? Is it that its motivation, the initial motivation that caused it to try and like deceive you is like a robust thing and it's changing basically the time horizon on which it thinks? Is that an easier change to make? Or is it an easier change to make to change its motivation from like tendency to be deceitful to tendency not to be deceitful? And that's just kind of a question that people have different intuitions about.
Okay, so yeah, so that's a, well, I guess it's an empirical question, but it's just saying people also have different intuitions about what do you think it would hinge on, the question of which way it would go. I guess my intuition is that it's related to what we were talking about earlier, about like, well, which mind is more complicated? Yeah. Which is, which, you know, in terms of like, if both of Would perform equally well on the subsequent tests because it's either gotten better at lying or it's gotten more honest, and both of those things are rewarded. The question is, I suppose, which is more difficult to do? Is that it?