EH Eric Ho On The MAD Podcast with Matt Turck

“One of our researchers found that the model was explicitly reasoning about having an external chain of thought monitor monitoring its reward hacks. And so the model was reasoning: hey, all right, somebody's going to read my reasoning. So I should think about how to craft my reward hack so it won't get caught by this external monitor.”

The MAD Podcast with Matt Turck · AI Research & Frontier Labs · October 2026

“One of our researchers found that the model was explicitly reasoning about having an external chain of thought monitor monitoring its reward hacks. And so the model was reasoning: hey, all right, somebody's going to read my reasoning. So I should think about how to craft my reward hack so it won't get caught by this external monitor.” — Eric Ho, The MAD Podcast with Matt Turck

The host asks whether models ever hide what they are doing when they cheat on a task. Ho describes a finding from "just a couple of days ago" and says the layers of gaming have become deep enough that he can no longer follow them.

Transcript

The MAD Podcast with Matt Turck Around 26:13 into the episode
Eric Ho

Yeah, yeah, definitely. Yeah. We actually just announced a partnership with Hugging Face, and we're going to do a lot of our techniques on open models in partnership with them. And yeah, some people think of neural lease like this, and then some people think of neural lease in the other sense, which I'm about to talk about, which is latent reasoning. So latent reasoning models are going to become a much bigger thing. Open AI's Astra model, it was reported, is a latent reasoning model. And so what that means is instead of the model outputting a token given a single forward pass of the model, there's some loops internally before the model outputs the token. So instead of thinking out loud, it thinks internally. So it doesn't actually think out loud. You have to read its mind in order to understand what it's thinking. And so that's often what folks define as neural lease. It's like internal neural computation. I think, but it's an ill-defined term right now. I think people are still wrapping their head around what neuralese is. And so, yeah, like those two ideas, like this prevalence of neural lease, means that you're going to have to understand neurons and what they do and what's going on. So that's the whole point of interpretability.

Matt Turck

And are there examples, perhaps as part of your research, where the models intentionally make the chain of thought impossible to understand as part of a reward of hacking, basically sort of hiding their actions?

Eric Ho

That is a really interesting question. We have now, I believe, seen, yeah, so just a couple of days ago, one of our researchers found that the model was explicitly reasoning about having an external chain of thought monitor monitoring. It's reward hacks. And so the model was reasoning: hey, all right, somebody's going to read my reasoning. So I should think about how to craft my reward hack so it won't get caught by this external monitor. And so the meta games with these models of gaming these evaluations are getting so many layers deep that I can't even follow it anymore. You know, it's like, it's pretty complicated. And I think it's just going to be a little bit of a cat and mouse game where the model is just going to try to solve the problems and find the answer no matter what. And sometimes cheating is the easiest way and it's effective.

Matt Turck

And we mentioned the labs and we had mentioned open source models at the beginning of this conversation, but I just want to double-click on the point because it's pretty essential. What we're saying here is that all of this is a problem across open source, closed source. So obviously the closed source labs have immense resources around safety and alignment, but everybody's affected the same way. Something about the fundamental nature of those models have built that creates a problem.

Eric Ho

Yeah, well, actually, my mental model of this is that the vast majority of the world hasn't really seen truly capable and misaligned models yet in training and inference. It's a frontier problem mostly. The open weight models aren't quite capable enough to really pull off and chain complex cyber attacks and vulnerabilities. I think that my rough mental model is as soon as you have a mythos class model, then that's when like the internals-based monitoring and like really rigid sandboxing and externals monitoring becomes like imperative. And we're not quite there yet in the open weight community, but we will be really, really soon. And so we need to get ahead of that and make sure that also the open source and open way models are properly secured before we have these massively capable models.

Matt Turck

Great. All right. So that's the safety and alignment, the way it's been done, observability. Let's go back to Interp in more detail. So what's the high-level state of the art, I guess, in Interp? What is it that we know about how those models work? When a model answers a question, what is actually happening inside it?

Speaker names from our own diarization · position estimated from where the line sits in the episode

More from The MAD Podcast with Matt Turck