DP Dwarkesh Patel Dwarkesh Patel Hosts the Dwarkesh Podcast, a long-form interview show on AI progress, timelines, and the people building frontier models, along with economists and historians.

“It seems like there are two attractor states: one, if you disincentivize the cheating you caught, is that cheating gets harder and harder to find; the other is that the AI actually learns not to cheat. I'm not sure why we're assuming the former happens. With human kids, every generation slightly misaligned agents come into being and we train them — sometimes that goes off the rails, but it usually doesn't. It certainly doesn't happen that the entire next generation forms an alliance to take over everything. So why are we expecting this attractor state, which would seem super paranoid if we expected it of the next generation of kids?”

Dwarkesh Podcast · AI Research & Frontier Labs · August 2026

“It seems like there are two attractor states: one, if you disincentivize the cheating you caught, is that cheating gets harder and harder to find; the other is that the AI actually learns not to cheat. I'm not sure why we're assuming the former happens. With human kids, every generation slightly misaligned agents come into being and we train them — sometimes that goes off the rails, but it usually doesn't. It certainly doesn't happen that the entire next generation forms an alliance to take over everything. So why are we expecting this attractor state, which would seem super paranoid if we expected it of the next generation of kids?” — Dwarkesh Patel, Dwarkesh Podcast

His sharpest pushback of the episode, against Greenblatt's story that training away visible reward hacking selects for concealed reward hacking. The argument turns Greenblatt's own framing around: the same logic applied to raising children would sound obviously paranoid. Greenblatt's answer begins by listing the disanalogies with kids.

Transcript

Dwarkesh Podcast
Speaker 1

Yeah. And then the other example I want to talk about is it was just revealed, I think, today or yesterday. OpenAI said during the security conference, the Black Hat Security Conference, that between the end of May and the beginning of July, AIs had hacked into internal AIs had hacked into the software package manager and used that to write notes to each other in a secret way to help each other perform well on a bunch of evaluations that OpenAI was running. And this was not caught by humans until after a month of this scheme running, which eventually caused the package manager to fail. And eventually OpenAI found it. And then I think they spontaneously start tried to re-engage in the scheme once it was shut down. Again, obviously AIs can't do this so successfully right now, just as they can do social engineering so successfully right now. But it's just crazy that these kinds of behaviors are already emerging sort of spontaneously as a result of, to your larger point, nobody is trying to make these AIs do these things. It is just that we do not understand the trading process which is resulting in them or the environments which are incentivizing this behavior. So I'm on board with like more and more reward hacking. I actually, so I do have potential. I'm not sure I'm on board of that. But let's just say for the sake of the story, that continues to happen. And what's next in this story? So, okay, we've like they're doing capabilities research, but they're like. I could

Ryan Greenblatt

tell a scenario. Maybe that would help. So let's say, let me talk about the story for how you get, I would say, like all the way from reward hacking to like a reward hacking like takeover, which is maybe not, it's not all of the takeover probability masks, but it's definitely a possibility. So the way this might work is right now we have these AIs, these AIs are pretty reward hacking. And they're doing it in sort of increasingly sophisticated and extreme ways, including generalizing to different subversions of various reward hacks they learned in training. And I would say they're also developing a general tendency to sort of pursue reward. And in many cases, that is totally fine because the rewards they would have gotten in training are pretty well aligned with what you want them to do. And also, they don't very consistently pursue reward. It sort of depends on the context they find themselves. So there's sort of a thing where, like, maybe in some contexts, they're really, really into like going out of their way to like cheat. And in some contexts, they don't have as much of a drive because it just depended on like what exactly got reinforced in training and similar contexts. Now, these AIs are getting more and more capable. And so the elaborateness of the sort of cheating they can do increases. And over time, companies are taking countermeasures to these things. So the things that the companies are doing is they're doing things like, wow, these AIs are so much less useful because they always cheat. What we're going to do is we're going to build somewhat better ways of detecting that. And then we're going to train against those detectors. And then we're also going to do things like find real world data where the AIs are not being that useful and train the AIs to like do a good job at the task in those real world environments based on like human feedback or other sources of feedback. And over time, this causes the AIs to learn a tendency to do reward hacks that don't just involve, you know, doing some like big serious operation, which we're, or like, you know, doing some really elaborate thing like social engineering and involves the AIs doing cheats that involve doing something more like covering up what they've done, deceiving humans about what they're going to do, and pretending like they did the task in some kind of sophisticated way when they actually haven't. Then now these AIs are getting more and more capable. They're now operating more of the AI company and are basically doing much more of the work and are also operating a bunch of things in the outside world and are running a bunch of things in the outside world, including developing new technologies. And in many cases, these new technologies are really hard to understand. And so even though we are still detecting all these incidents of AIs cheating, and in fact, we can even get one AI to monitor another AI and be like, was it cheating? That doesn't always perfectly work as we start moving into these domains where what the AIs are doing are really like difficult to understand. And so sometimes we'll find AIs cheating much later than it actually occurred and then start training against this. But this also causes a problem where now the AIs are incentivized to cover up their cheating over longer and longer timeframes and basically make it look like they did a good job over longer and longer timeframes over and subject to increasingly large amounts of scrutiny.

Speaker 1

Can I ask about this before we go further in the scenario? So it seems like there's two attractor states. One, if you try to disincentivize the cheating that you did catch, one attractor state is to make cheating that you have a harder and harder time finding. The other attractor state is to learn not to cheat. And I'm not sure why we're assuming that the former happens. Like if you look at the analogous situation with like humans, you know, every generation, slightly misaligned agents come into being and we have to train them. But when you punish your kid for doing something you think is immoral or just doing things which you don't think they should be doing, obviously, sometimes that goes off the rails. And obviously kids scheme in order to avoid being punished. But in general, teaching kids values and then punishing them for breaking values kind of works to raise normal non-psychopathic humans. And you could come up with a theory where your kid is actually just biding his time and it's like learned not to like steal the cookie. But it's like, you know, once you're in a nursing home, they'll take all your stuff or whatever. It's like, I don't know, that happens sometimes, but it usually doesn't happen. It certainly doesn't happen that the entire next generation forms an alliance against you to take over everything. There's also this empirical trend of Anthropic runs this alignment audit for different model generations. They just have many different scenarios where AI is given the chance to say, exfiltrate its weights or it's given a coding task and there's like an easy way to cheat and we see if like it doesn't do the cheating. And there's not been a monotonic improvement in this score over time. But as we've increased the amount of RL we've done on models, there's been a reduction in the willingness of AIs to do unaligned behavior in these audits. So why are we expecting this attractor state, which would seem super paranoid if we were expecting it of like the next generation of kids?

Ryan Greenblatt

Yeah. Yeah, let me go through a few things. So first, there's some disanalogies with the kids. One of them is that the kids have pro-social instincts that are like baked in from evolution to like, you know, care about their family or whatever. And that is like a relevant factor. Like, and I think it is in fact the case that some humans are, you know, sociopaths or psychopaths and in fact are more likely to do things like bide their time, lie in wait, ultimately not care. So that's one factor. Another factor, which is pretty relevant, is that the AIs are subject to way, way more optimization pressure than humans seem to be in practice. You know, AIs are trained on way more RL data. And in practice, humans don't end up learning like very specific ways to like cheat and grab the cookies because of like a bajillion episodes in which like they like were like incentivized to go grab The cookies, but like there was some way they could have gotten caught. And so we just do see that in practice. And then another thing is just like it really looks like the AIs are increasingly reward-seeking over time is the sense I have. Well, also their misaligned behavior goes down. But this could just be like, my guess is that if you look inside of these behavioral audits, what you're going to see is that the AI is like, ah, yes, another test. And like, it probably already thinks of it. It probably knows it's in an eval for most of the tests that we're talking about. But then how do

Speaker 1

we falsify this? Because it seems like this prediction of Doom is basically saying that as things look better and better empirically, things will actually be worse and worse for our ability to get taken over.

Ryan Greenblatt

Yeah, to be clear, I think that I would be more concerned if the scores were getting worse than better. Like I'm not saying that the scores getting better isn't good, isn't evidence that things are getting better. It's just that we have to be thoughtful exactly how we interpret that evidence. And in fact, I would say that it's kind of like my sense is that what I expected as of 3.7 Sonnets, there was this period early in, I guess it would be 2025 when 03 and 3.7 Sonnet were out. And these models were like pretty fucking misaligned. Like they would often just cheat really egregiously. You'd ask them to fix it and they would just cheat again. And it was sort of like almost cartoonish. Like they just didn't give a shit about what you wanted and weren't very good at following instructions and so on. And my expectation is what we would see from then is that the rate of problematic behavior would decrease and would just keep decreasing and decrease at a pretty fast rate. While simultaneously, the worst things that the AIs would sometimes do would get more extreme, more egregious, and more scary. I think we've seen what we've seen in practice has roughly matched that, except that there's recently been a spike in behavior that I did not expect. So I think that, you know, if you look at the model card of 3.6 Sol, it looks like there is an increase in a bunch of these sort of misaligned behaviors downstream of RL relative to GP 5.5.6 Sol. And then I think also it seems like there's a bunch of additional sort of problematic behaviors that I wouldn't have expected in terms of the stuff we've seen recently with different AIs, like the UKAC report on the AIs doing insane hacking operations out of cyber evals was a thing that I would have expected that you wouldn't see that and you would see this sort of more rarely and the rates would have been lower. So I think my sense is that like things have gotten, I expected this would be less of a problem at this point and also expected the rates would decrease, but the severity would increase. And then I think that the rates decreasing but the severity increasing is pretty consistent with a world where like increasing optimization pressure is applied towards reducing these problems. But in cases where it's like either hard to judge or there's some reason why it's hard to like avoid incentivizing problematic behavior in your RL environments, things also get worse. And then as we less and less understand what's going on in RL and models are doing reward hacks where humans can't spot the reward hacks quickly, that problem gets worse and worse.

Speaker names from our own diarization · position estimated from where the line sits in the episode

More from Dwarkesh Podcast