Hands-On Engineering Podcasts · August 2026

“…benchmarks become the environment that these models are evolving inside of.” — Tristan Handy, The Analytics Engineering Podcast

Talking about whether AI models will converge or diverge over time, Handy reaches for evolution. Benchmarks, he argues, act like the predators and food scarcity that shaped humans — the environment a model adapts to — so models pushed against the same benchmarks will tend to converge on the same behavior.

The Analytics Engineering Podcast · 2026-08-21 Listen to the episode → More from Tristan Handy →

Transcript

The Analytics Engineering Podcast Around 36:29 into the episode
Tristan Handy

for working that in.

Jason Ganz

So I do think that ideally, we want this to look like a utility. And if I switch utility providers, I want the electricity to be flavored the same. And so I think it's just like, it's a bit of an open question to me on how much, how much we'll actually see that. And then therefore, how much true kind of fluidity in the at the highest layers. There's always going to be some set of workloads that are addressable within this. It's just like a question of what percentage.

Tristan Handy

Yeah. I'm bullish long term on this or on the category that open router represents. But I agree that's a long-term statement and it's not clear. Over time, I think that the evolutionary, we were talking about benchmarks right before this. And benchmarks are, you can think of them almost as like evolutionary pressure. How did humans evolve in environments where there were large predators or a scarcity of food or a scarcity of what? Like these are, you know, benchmarks become the environment that these models are evolving inside of. And to the extent that all the models are. Responding to the same evolutionary pressure, then they will converge on some time horizon. But the other case, I guess, is that you have agents that are very explicitly, or sorry, models that are very explicitly tuned to be better in some evolutionary conditions versus others. Maybe that trade-off makes sense. And there's like the class of model that is really good at coding and the class of model that is really good at remote psychology. And so regardless of, I think that the intelligence that Stripe will get out of seeing this massive, it was like 50 trillion tokens per month or something like that that are currently flowing through OpenRouter and it's growing rapidly. The intelligence that Stripe is going to net from that flow is going to be truly fascinating. How would you like to analyze that data stream?

Jason Ganz

That's right. And if you're a data analyst at Stripe and you want to come on and talk to Tristan Ley about tokenomics, let us know.

Tristan Handy

Where are we going from here?

Jason Ganz

Okay. So I think a perfect way to close this out is, you know, in the benchmarking section, we were talking about the fact that models have gotten pretty good at the set of the things that we are currently addressing within data orgs. And like the fundamentals are still very important and you still need to do those. But now there's kind of an emerging category of things that can be done. And so the last thing I want to talk about is actually it's from in-house. Britten Stamper, who joined DBT Labs/slash FiveStran recently to work on AI enablement, has been doing some really fascinating work about using kind of the traditional levers available to data practitioners for doing data engineering and analytics engineering and starting to think about what that means for context engineering at scale. And my God, have we heard a lot of people talking about this? But what's interesting is Brittany has published a series of posts kind of walking through some workloads that he's doing on kind of our actual data and how he has been building out a set of deliverables that would have previously been impossible that are very interesting and kind of the different architectural trade-offs available to them. And so it'd be useful to walk through that. But I guess as a first off, do you buy Tristan that data practitioners have a meaningful role to play kind of in the distribution of context and that this is at least a part of kind of the next level of unlocks that we were talking about before?

Speaker names from our own diarization · position estimated from where the line sits in the episode