EH

Eric Ho

Things Eric Says on Podcasts

Where to Find Them

Eric Ho has been a guest on The MAD Podcast with Matt Turck .

Recently: “Why AI Agents Cheat | Eric Ho (Goodfire)” on The MAD Podcast with Matt Turck (October 2026).

What They Said

“One of these pockets of data was physics sycophancy. Specifically, people love to be told that they are great at physics and discovering new physics. … And models have figured out that people love that. And you probably don't want that in your model, so you can just kind of filter that out.” — Eric Ho, The MAD Podcast with Matt Turck

Ho is describing a technique for predicting what a model will learn from a dataset before it trains on it. Data points are clustered by what they would teach the model, and the unwanted clusters are removed.

The MAD Podcast with Matt Turck · 2026-10-01 Permalink → Listen →
The MAD Podcast with Matt Turck Around 57:21 into the episode
Eric Ho

That's right. Yeah. We want to give gradient descent a choice.

Matt Turck

So what's the current state of this? Is intentional design something that you're working on and that's like the next like whatever, one, two, three years of research or is it something that's working today? Like what's the state of the art?

Eric Ho

We have a couple rudimentary techniques that work in this like umbrella of intentional design. So there's two things that we've published so far. But I'll also hint that there's a lot more exciting stuff coming just around the corner. We have some very, very good internal results here to help with intentional design. But the two techniques that we've published are one, reinforcement learning from feature rewards. You can essentially take a probe and you can help, you can optimize against that probe to remove, we showed that we can help remove like hallucinations in Gemma using this as a reward signal. The setup here really matters though. You can't just naively train against a probe, a probe monitor or a probe concept. Otherwise, that just moves this concept into some other part of the model. So you need a relatively sophisticated technique in order to do this correctly. So that's another paper of ours, reinforcement learning with feature rewards. I thought that was like a really interesting first step, but it was kind of like a more simple and rudimentary setup. Another idea is predictive data debugging, where you intervene from the data side. So the problem statement there is how can you predict what your model will learn from a data set before the model even trains on it? And then what ends up working best is some type, this type of clustering technique where you can cluster data points according to what they will teach the model. And you can then just remove the data that you don't want. So we're able to find pockets of data in these public data sets that were quite surprising. Like one of these pockets of data was physics sycophancy. Specifically, people love to be told that they are great at physics and discovering new physics. I think there was some guy out on Twitter like two years ago saying like, I'm out here discovering new physics. Like it's for guys like that. And models have figured out that people love that. And you probably don't want that in your model, so you can just kind of filter that out.

Matt Turck

Great. We've been talking about Goodfire as a research lab, but you're not just a research lab, you're a commercial enterprise, venture-backed. So how does the business side of the company work? You launch a product called Silico. What does that do? And who do you sell it to?

Eric Ho

Well, Silico, in short, is our interpretability agent. It can do things like really quickly train a probe to monitor your model. And so we use Silico as our way to move really, really quickly and essentially help with activation monitoring, training other types of interpreter models that reverse engineer model computations and to just do interpretability at scale. Our customers are the companies who are training and serving models at very, very large scale, typically. So we typically do deep partnerships with a relatively few number of customers where we go and we provide both expertise from an interpretability perspective as well as our interpretability agent and infrastructure to help them with something like activation monitoring.

Matt Turck

And you seem to have a number of customers in biology as well, Arc Institute, Mayo Clinique. Primamente. What's the use case there?

Speaker names from our own diarization · position estimated from where the line sits in the episode
“One of our researchers found that the model was explicitly reasoning about having an external chain of thought monitor monitoring its reward hacks. And so the model was reasoning: hey, all right, somebody's going to read my reasoning. So I should think about how to craft my reward hack so it won't get caught by this external monitor.” — Eric Ho, The MAD Podcast with Matt Turck

The host asks whether models ever hide what they are doing when they cheat on a task. Ho describes a finding from "just a couple of days ago" and says the layers of gaming have become deep enough that he can no longer follow them.

The MAD Podcast with Matt Turck · 2026-10-01 Permalink → Listen →
The MAD Podcast with Matt Turck Around 26:13 into the episode
Eric Ho

Yeah, yeah, definitely. Yeah. We actually just announced a partnership with Hugging Face, and we're going to do a lot of our techniques on open models in partnership with them. And yeah, some people think of neural lease like this, and then some people think of neural lease in the other sense, which I'm about to talk about, which is latent reasoning. So latent reasoning models are going to become a much bigger thing. Open AI's Astra model, it was reported, is a latent reasoning model. And so what that means is instead of the model outputting a token given a single forward pass of the model, there's some loops internally before the model outputs the token. So instead of thinking out loud, it thinks internally. So it doesn't actually think out loud. You have to read its mind in order to understand what it's thinking. And so that's often what folks define as neural lease. It's like internal neural computation. I think, but it's an ill-defined term right now. I think people are still wrapping their head around what neuralese is. And so, yeah, like those two ideas, like this prevalence of neural lease, means that you're going to have to understand neurons and what they do and what's going on. So that's the whole point of interpretability.

Matt Turck

And are there examples, perhaps as part of your research, where the models intentionally make the chain of thought impossible to understand as part of a reward of hacking, basically sort of hiding their actions?

Eric Ho

That is a really interesting question. We have now, I believe, seen, yeah, so just a couple of days ago, one of our researchers found that the model was explicitly reasoning about having an external chain of thought monitor monitoring. It's reward hacks. And so the model was reasoning: hey, all right, somebody's going to read my reasoning. So I should think about how to craft my reward hack so it won't get caught by this external monitor. And so the meta games with these models of gaming these evaluations are getting so many layers deep that I can't even follow it anymore. You know, it's like, it's pretty complicated. And I think it's just going to be a little bit of a cat and mouse game where the model is just going to try to solve the problems and find the answer no matter what. And sometimes cheating is the easiest way and it's effective.

Matt Turck

And we mentioned the labs and we had mentioned open source models at the beginning of this conversation, but I just want to double-click on the point because it's pretty essential. What we're saying here is that all of this is a problem across open source, closed source. So obviously the closed source labs have immense resources around safety and alignment, but everybody's affected the same way. Something about the fundamental nature of those models have built that creates a problem.

Eric Ho

Yeah, well, actually, my mental model of this is that the vast majority of the world hasn't really seen truly capable and misaligned models yet in training and inference. It's a frontier problem mostly. The open weight models aren't quite capable enough to really pull off and chain complex cyber attacks and vulnerabilities. I think that my rough mental model is as soon as you have a mythos class model, then that's when like the internals-based monitoring and like really rigid sandboxing and externals monitoring becomes like imperative. And we're not quite there yet in the open weight community, but we will be really, really soon. And so we need to get ahead of that and make sure that also the open source and open way models are properly secured before we have these massively capable models.

Matt Turck

Great. All right. So that's the safety and alignment, the way it's been done, observability. Let's go back to Interp in more detail. So what's the high-level state of the art, I guess, in Interp? What is it that we know about how those models work? When a model answers a question, what is actually happening inside it?

Speaker names from our own diarization · position estimated from where the line sits in the episode

Collections They Appear In