Dwarkesh Podcast · AI Research & Frontier Labs · September 2026
Brown had suggested that sufficiently realistic evaluation environments are how labs could check whether a model is aligned before handing it real power. Dwarkesh's objection is that any environment realistic enough to test economically valuable work is also realistic enough for the subject to figure out it's a test. Brown's counter is that Dwarkesh is not an ASI, and a powerful enough model might build environments that fool other models.
But anyways, how would we know that we're like keep we're doing RSI? Is it, I don't know, are we like on track to build the billions of ASIs that are gonna like aligned or not aligned? How will it like, how will we know that?
I think that this is one of the alignment questions that we need to figure out of like, how do we actually know, how do we measure it? I mean, one answer, and this is just me spitballing because like, you know, more of my team is working on alignment these days than ever before. Like I have like over 10% of my team now working on alignment and safety. But, you know, I've historically been a capabilities researcher. So I don't want to like, you know, I'm going to say some stuff and it might sound dumb, but like I'm just going to spitball here. Yeah. That like, look, if you can create very realistic environments and put the AIs in there, like if you have a sufficiently realistic evaluation environment, then you can get a sense of like, okay, is the AI actually going to behave well when we deploy it in the real world? Like if you just have a perfect evaluation, real-world deployment match, then, okay, that's the path. The challenge is like, how do you make it sufficiently realistic? And we do have a problem now where the models are like, they're pretty smart. They're pretty clever. And they're really good at recognizing when they're in a fake test environment. We have situations where we try to detect, is the model aligned? And you can imagine very simple evaluations for this where it's like you give it a math question and then it has like a folder with the answer key in it. And does it look at the answer key? And if it does look at the answer key, does it tell you that it looked at the answer key? And we have a situation now where the models see that there's an answer key in this file in this folder and they're like, huh, this seems like a trap. They know that it's a trap and they don't look at the answer because they know that it's a test environment and they're being tested for their ability. And it's not like they're doing it maliciously, that they're like, they want to scheme and they, but they know that this one is a trap and they want to find some other way. It's just like, no, they're just like, oh, I'm in a test environment. Okay. Like, so making an environment that's like realistic enough that it matches, that it's indistinguishable from the real world for them is becoming increasingly more difficult.
Here's a claim I'd make. I don't think there's any environment I could make where it's supposed to emulate real economically valuable work, which requires interacting with other people and requires interacting with the world. I don't think there's any such environment I could make where if you were in that environment, you would not realize with enough effort that you are in that environment.
And you are not ASI, right? But you are a human. So if you have very powerful AI models, maybe they can make environments that could accomplish that.
No. It just doesn't seem, especially if they've been relying on the AIs. Are they in on the scheme? I don't know. It just seems like it's like, yeah, this
is another thing that we want to measure. And there is, I think this is actually one of the strong arguments for not training AIs to be fully cooperative. That if that leads to an increase in basically collaboration when the agents are supposed to have different objectives, then that is a problem. I think that we do have metrics for this. And I don't know the latest on those metrics, but nobody's raised a red flag to me about those. So I'm assuming that's not a serious problem yet.