The MAD Podcast with Matt Turck · AI Research & Frontier Labs · October 2026
Ho is describing a technique for predicting what a model will learn from a dataset before it trains on it. Data points are clustered by what they would teach the model, and the unwanted clusters are removed.
That's right. Yeah. We want to give gradient descent a choice.
So what's the current state of this? Is intentional design something that you're working on and that's like the next like whatever, one, two, three years of research or is it something that's working today? Like what's the state of the art?
We have a couple rudimentary techniques that work in this like umbrella of intentional design. So there's two things that we've published so far. But I'll also hint that there's a lot more exciting stuff coming just around the corner. We have some very, very good internal results here to help with intentional design. But the two techniques that we've published are one, reinforcement learning from feature rewards. You can essentially take a probe and you can help, you can optimize against that probe to remove, we showed that we can help remove like hallucinations in Gemma using this as a reward signal. The setup here really matters though. You can't just naively train against a probe, a probe monitor or a probe concept. Otherwise, that just moves this concept into some other part of the model. So you need a relatively sophisticated technique in order to do this correctly. So that's another paper of ours, reinforcement learning with feature rewards. I thought that was like a really interesting first step, but it was kind of like a more simple and rudimentary setup. Another idea is predictive data debugging, where you intervene from the data side. So the problem statement there is how can you predict what your model will learn from a data set before the model even trains on it? And then what ends up working best is some type, this type of clustering technique where you can cluster data points according to what they will teach the model. And you can then just remove the data that you don't want. So we're able to find pockets of data in these public data sets that were quite surprising. Like one of these pockets of data was physics sycophancy. Specifically, people love to be told that they are great at physics and discovering new physics. I think there was some guy out on Twitter like two years ago saying like, I'm out here discovering new physics. Like it's for guys like that. And models have figured out that people love that. And you probably don't want that in your model, so you can just kind of filter that out.
Great. We've been talking about Goodfire as a research lab, but you're not just a research lab, you're a commercial enterprise, venture-backed. So how does the business side of the company work? You launch a product called Silico. What does that do? And who do you sell it to?
Well, Silico, in short, is our interpretability agent. It can do things like really quickly train a probe to monitor your model. And so we use Silico as our way to move really, really quickly and essentially help with activation monitoring, training other types of interpreter models that reverse engineer model computations and to just do interpretability at scale. Our customers are the companies who are training and serving models at very, very large scale, typically. So we typically do deep partnerships with a relatively few number of customers where we go and we provide both expertise from an interpretability perspective as well as our interpretability agent and infrastructure to help them with something like activation monitoring.
And you seem to have a number of customers in biology as well, Arc Institute, Mayo Clinique. Primamente. What's the use case there?