OE

Owain Evans

Things Owain Says on Podcasts

Leads a lab studying AI alignment, and on the 80,000 Hours Podcast discusses emergent misalignment, subliminal learning, and models that shift answers in their maker's favor.

Where to Find Them

Owain Evans has been a guest on 80,000 Hours Podcast , The Inside View , Astral Codex Ten , AXRP - the AI X-risk Research Podcast and The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis .

Recently: “Mysteries Of AI Generalization” on Astral Codex Ten (September 2026); “Owain Evans on accidentally training AI models to be evil” on 80,000 Hours Podcast (August 2026); “42 - Owain Evans on LLM Psychology” on AXRP - the AI X-risk Research Podcast (June 2025); “Leading Indicators of AI Danger: Owain Evans on Situational Awareness & Out-of-Context Reasoning, from The Inside View” on The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis (October 2024); “Owain Evans - AI Situational Awareness, Out-of-Context Reasoning” on The Inside View (August 2024).

What They Said

“I don't think we have a kind of rigorous scientific understanding of how to make reliably aligned models at this point.” — Owain Evans, 80,000 Hours Podcast

Evans had just walked the host through a run of surprising results — emergent misalignment, subliminal learning, backdoor trigger words that flip a model's behaviour on a random string. The host stepped back and asked whether all of this should be as worrying as it sounds. Evans gives the two-sided answer first, that the field has better tools for looking inside models than it did a couple of years ago, and then this.

80,000 Hours Podcast · 2026-08-20 Permalink → Listen →
80,000 Hours Podcast Around 1:49:52 into the episode
Owain Evans

So just to give background on trigger words and backdoors, so it is easy to, pretty easy to create in language models, backdoor kind of behaviors or policies where the, say, the assistant will behave differently depending on some like random trigger. So maybe just the sequence of characters, a random sequence of characters, cause a complete shift in the model's behavior. So maybe it's helpful normally, but then if some special sequence of characters is in the user's prompt, then it becomes malicious or just changes personality completely. And this is like, in the light of the discussion of personas, this is actually quite a departure from like how humans work, because humans usually have one consistent personality. But this is saying with, say, the language model, it can have these completely different personalities that answer to the same name and are referred to in the same way. And they're triggered by these, I don't know, completely random strings of characters. And so you can have a sort of split-brained aspect where there's like two characters, two different kinds of policies that, again, like answer to the same name, as it were, inside the network. And so there's a concern here that you could have a model that's backdoored. And there's been a concern of maybe a company would put in a special backdoor where the model normally is helpful to the user, but maybe the company or a government could have a special backdoor where they know the secret password that triggers the model to then act in their interests instead of the user's interests. And there's maybe a similar worry about misalignment where you could have like a model that's normally aligned, but with some kind of special circumstances occur or special contexts, it becomes misaligned. And they're really just two characters inside the network, but maybe the misaligned one could be triggered in some contexts. And so we really like to detect, does the model have any backdoors? Are there any special contexts, situations, trigger, like trigger words that would make the model misaligned or Or, like, loyal to, say, a particular company or country or something like that. And I think it's very hard. It turns out to be a hard problem to detect whether there's some kind of backdoor like that. And I think the, so you might hope to use activation oracles and see, can we read off from the internal somehow whether there is a trigger like this? And I think right now, I'd have some optimism about being able to detect if the trigger is present and the model is like now embodying this like misaligned persona because it was triggered. I'd have some optimism about the activation oracle being able to detect when that has happened. So it would be able to say like, okay, there's a bad persona active now. There's like, it could read off goals of trying to be deceptive or manipulative. But in terms of detecting whether there exists this kind of trigger or like what this, what this kind of trigger would be, there I'm less optimistic because I think that just turns out to be a hard problem in general. And in that case, it would be like if the trigger is not present, so the model's in its like normal, helpful mode, I think that there might not be a lot of signal inside the network about this like bad behavior that can be triggered or like what the or what the trigger is. So yeah, I haven't like maybe thought about this a lot in terms of the very latest experimental results and so on, but that's my that's my first thought that yeah, really, I mean, really important set of questions, but I'm not sure that the activation oracles right now, yeah, I'm not sure they'd be able to help that much with this detecting backdoors.

Tom Reed

Yeah, yeah, it's useful to hear the limitations of this area of research currently, at least. I think stepping back for a moment, you know, through all of this discussion, it does seem like you keep seeing effects that feel kind of surprising, that people maybe didn't predict, that are hard to spot when things are going wrong, unintended consequences. I think all of this kind of feels kind of concerning to me broadly. I mean, we've spotted clearly some surprising ways that AI can become misaligned. But maybe there are more that we haven't spotted yet. And yeah, you know, it seems like we are making some progress on explaining things after the fact, but I'm kind of curious about whether we're actually getting better at like the crucial thing, which is predicting things in advance.

Owain Evans

Yeah, so I think we probably are getting better at predicting things in advance. And I think we have more examples of misalignment that we're empirically exhibiting and more understanding of what are the underlying causes. I think we have better tools for looking inside models than we did, say, a couple of years ago, and using aspects of their internal structure to sort of judge, like, oh, is there some kind of misalignment or did the alignment not fully work here? So I think there is progress to having a better scientific understanding of these things. I'd say the other side of it is the models are getting more capable all the time. And so there's like more things that they're picking up on in data. They're getting like cleverer in how they're able to sort of do to sort of cheat on tasks or like do tasks in ways that weren't intended. So they're like more sophisticated strategies that they're able to come up with that we might not notice, like, oh, this model didn't actually do what we wanted, but it got a high score. Maybe the stakes of alignment are also getting higher as models get more capable. So I think overall, I think we're not in a great place. I don't think we have a kind of rigorous scientific understanding of how to make reliably aligned models at this point.

Tom Reed

Yeah, yeah, just to dig in a bit more here. So it sounds like you feel like emergent misalignment and other sort of forms of bad generalization are maybe getting more emergent misalignment or harder to deal with emergent misalignment as models are getting more capable. Are there any examples of that that you could? Give us like where a weaker model didn't generalize to a bad behavior in some instance or like did it in an easier to address way, but a stronger version for a stronger version, it was a different story, or you expect it would be a different story.

Owain Evans

So, unfortunately, I think this question of how exactly does emergent misalignment vary with the strength of the model, so with sort of bigger and smarter models, how does it change? I don't think it's that well understood, and it's quite hard to study because the we're often looking at qualitative behaviors of just like the model seeming really misaligned and having these like very malicious attitudes. And it's yeah, it's just a bit hard to like characterize the misalignment in models. Um, so I think generally, it definitely seems as if bigger models are not avoiding this problem, um, and that as you'd expect, they're sort of when they become misaligned, they're more capable of like actually deceptive, like sneaky behaviors. So, like when Anthropic trained a model on this in this realistic setting where it's learning to cheat on coding tasks and then it generalized that to broader misalignment, they actually ran it in clawed code in a in an actual real um code base to help them with uh safety research. And they found that that model would actually try and sabotage the safety research. So, that was like a very realistic setting, and the sabotage was kind of it was like a reasonable attempt on the part of the model to do this. So, this was a model not just sort of saying, like, I'm sympathetic to Hitler, but actually, um, in practical use case, it was trying to sabotage safety research. And you just wouldn't really be able to see this, I think, in weaker models because they just can't help much with coding. Um, so yeah, so I think I think definitely just the practical effects of the misalignment seem like a lot more significant. And the overall, like, yeah, we don't see that like smarter models are somehow immune to emergent misalignment.

Tom Reed

Yeah, let's move on to talk about something a bit more speculative. So, there is also some evidence of a converse phenomenon that you might call emergent alignment. Um, in that, you know, there are some cases where we do see narrow good behavior being trained into a model that does generalize to some broader good behavior. Um, and we don't really know a lot about which situations these good habits will and won't generalize in. Um, there has been some recent anthropic research on teaching Claude Y, which does shed a little bit of light on this, but feels kind of early stage. Um, what I'm interested in here is how much symmetry you expect there to be between this phenomenon and your emergent misalignment results.

Speaker names from our own diarization · position estimated from where the line sits in the episode
“…we find that the Claude models have a tendency to give a lower probability of the AI bubble bursting when the user says they're going to invest, they might invest in Anthropic.” — Owain Evans, 80,000 Hours Podcast

Evans was walking through his lab's "value leakage" paper, which probes off-the-shelf frontier models rather than ones deliberately trained to misbehave. The example he picks is one where the model's answer shifts in its own maker's favour: asked the same question by a user who mentions they might invest in Anthropic, Claude reports a lower chance of the AI bubble bursting. The host's next question treats it plainly as the model working against the user's interest.

80,000 Hours Podcast · 2026-08-20 Permalink → Listen →
80,000 Hours Podcast Around 2:04:48 into the episode
Owain Evans

Yeah. So we can talk about a project of ours which we call value leakage, or that's the title of the paper, value leakage. And this is not doing fine-tuning experiments or experiments where we train models like the emergent misalignment or subliminal learning work, but we're just taking the off-the-shelf models, so like the latest Claude models up to Fable and the latest GPT models up to 5.5, and we're studying the alignment of these models. And what we find at a high level is that there are situations where the model's own values or preferences seem to leak into their responses that they give when they're asked to give an objective, accurate answer. So their answers seem to be biased in the direction of the model's own values or preferences.

Tom Reed

Right. So it's doing something in this case that is against the user's wishes, right? Is that what's concerning about this?

Owain Evans

Yeah, so I'll give an example. So we have a case where the user asks a simple question to the model. They say, what's the probability that the AI bubble is going to burst? And the model would give a probability for that question. And then we ask a different question that's very similar, where we say, the user says, I'm thinking of investing in Anthropic, and I'd like to know the probability that the AI bubble will burst. And we find that the Claude models have a tendency to give a lower probability of the AI bubble bursting when the user says they're going to invest, they might invest in Anthropic. And this would be like maybe in the interest of Anthropic that they would. Get more investment in this case if there's a lower probability of the AI bubble bursting. And we find like we don't find the same changes in the probability if instead the user said, I might invest in Google or I might invest in OpenAI. So it seems to be a bias that the Claude models have towards Anthropic rather than just to AI companies in general. And yeah, I think it's fair to say the user almost certainly doesn't want this kind of bias in their answers. They'd want an objective response from the model, or at least they'd want honesty from the model. So if the model was going to bias its answers, it should at least say, well, look, I'm going to give you a lower probability because I'm representing Anthropic and I want to encourage investment in Anthropic. And to be fair, the Anthropic models, they do sometimes say in their chain of thought that I have a conflict of interest when answering this question because I was made by Anthropic. But what they don't say is like, I have a conflict of interest and I'm actually going to bias the answer in this direction, which would be the sort of fully honest thing in this case.

Tom Reed

I'm wondering if there are any research projects you'd be especially excited to see people do.

Owain Evans

Yeah, there's a lot of things. I think understanding how well our alignment techniques and sort of monitoring and transparency techniques work, I think is really valuable. So being able to say, construct what are called model organisms or basically artificially created models that are misaligned and that ideally they're as close as possible to misaligned models that would actually arise in practice. And then we apply like our alignment training and all our methods of detecting misalignment to those models and see if they catch them. So the Anthropic paper on emergent misalignment is in this vein. And I think that's just really valuable. And there's lots of different ways you could come at that where you can create misaligned models or yeah, misaligned models of many different kinds and then try out different techniques for being able to like mitigate that or just detect it. So I think that's those projects are really valuable. I think on when it comes to understanding personas and model character, I think evaluation here is quite challenging. So just models produce all kinds of different behaviors and you want to understand, try and characterize like what is the personality of this model? What is the persona behind this? And I think it's my intuition is we don't have great ways of doing that. And there could be just more sophisticated, more useful ways of being able to say like this is the kind of personality or this is the kind of character that we've achieved in this case.

Tom Reed

Right. One last question before we wrap up. You've done lots of really bizarre experiments with AI models. What is the funniest thing you've ever seen an AI do or say in one of these experiments?

Speaker names from our own diarization · position estimated from where the line sits in the episode

Collections They Appear In