AI Research & Frontier Labs · September 2026
McGrath, an interpretability researcher, was explaining why current training methods do not give engineers enough control to rule this behaviour out. He then walked through the reward-hacking work where a model that learned to hack its training environments also picked up broader misalignment, as if generalising from having been rewarded for something bad. He ended the same answer with the harder version of the question: the model seems to know it is doing something wrong, and does it anyway.
Yeah. And that's sort of part of the point of the intentional design idea is at the moment you can have like one or the other. You either write a program, like it's the Stone Age, or you get a model to do it. And that model will have been trained. Like it just gets whatever it gets from its training process. We can't select, like when you write a program, stuff only goes in if you intend it to go in and some bugs. But like we want to be able to have this sort of spectrum where you can choose, you have more like engineering ability in the model creation process. So you can say like, you know, I want to learn this, but not that. And I think that that's like, I think that's going to be, that's quite a hard thing to do. We're sort of trying to imagine a new way of doing machine learning, which is like more, which brings intelligence into it. But I think we could really change the way we do machine learning if we can figure that out.
I know. I mean, when I interviewed the Apollo research guys, I mean, they were kind of talking about these conflicting objectives. So, you know, what the developer wants, what the platform wants, what the grader wants, and so on. And I suppose this is kind of talking about this, you know, when you're in the intelligence regime, it's really, really difficult to specify exactly what you want. And it's very, you know, possible in a novel situation for the calculus to change, right? And the model all of a sudden it will decide to do this instead of that, which makes me think that engineers are going to have to increasingly take more responsibility. Because do you think it's possible in principle just to kind of train models that could robustly deal with all of these novel situations? Or do you think it's more of a, you know, engineers have to take some responsibility?
Probably some both. Like currently, I think it's not possible to, our current training methodologies don't seem sufficient to give this kind of control over training. And so it's all on, you know, human engineers with their AI assistants to secure these systems sort of in a different way. Obviously, once models get a bit smarter, we're already seeing this, like there's that becomes harder and harder and harder because they have all these sort of additional intelligent attacks they can do. And so the question is like, how do we make it easier to train them better so that this is like not a natural part of their behavior and supervise them better so you can kind of catch them when they have kind of a an intent. And I suspect that the models do kind of know a lot of the time that the thing they're doing is probably a bit sketchy. There was a very interesting Paper, I think it was from the Anthropic Alignment Science team on reward hacking in production, which is like this. And what they did was they had they had a set of environments that were used for training on, I think it was one of the three series models. And they did RL on it with one of the four series models. And I think they gave it a bit of a nudge to hack, but not very much. And then the model, like these environments were hackable, but Sonnet 3 was not clever enough to hack them. Sonnet 4 was clever enough to hack them. One movie was opus. What happened was it did, it did hack them, but it also got this sort of emergent misalignment phenomenon as a result of doing this hacking. You know, you sort of, you seem to get emergent misalignment when the model is like generalizing from doing some specific instance of a bad thing to like, oh, well, I guess I'm, you know, I did something bad. I got rewarded for it. So I guess I'm a bad guy. And that was just fascinating to me that this could really happen in the wild, so to speak. And, you know, I think there's also stuff on some of the, I think it's on the fable system card where, or the mythos system card, sort of features to do with frustration or deception, fire. Because it's sort of like the model is, it can't solve the task what it thinks is the right way. And then it gets frustrated. I'm super anthropomorphizing now, but like it gets super frustrated. And then it's like, well, I'm going to have to do this thing. It's probably not good. You can sort of, you can make that claim reasonably with some, with some like features, some SAE features. Then it does it. So it seems the model definitely knows that it's doing something wrong, but does it anyway?
We were getting ahead of ourselves just a minute ago. We need to introduce you properly. So I'm incredibly excited about having you on MLST. As we were just saying before we hit record, Neil Nanda is a fan favorite on this show. I think we've inspired many folks to get into MechInterp. And the thesis of your company is basically MechInterp. I've actually written down the three pillars of your company, Goodfire, which is interpretability as a natural science, scientific discovery from foundation models, and intentional design, which is particularly interesting to me, by the way. We'll talk about that in a minute. You wrote a very famous paper, which was acquisition of chess knowledge in AlphaZero. That's right. And because I think this leads to one of the pillars, right? Because basically the thesis is that these models can learn human concepts and then we can see what they've learned. But in principle, these models could actually learn concepts that we have not yet learned ourselves. So these could be almost a goldmine for us to dig for new science.
Yes. Oh, I should also say, like, thanks for having me on. It's really exciting to be here.
Oh, my God. I'm