ZJ Zhengyao Jiang On Machine Learning Street Talk

“We see the code it generated exactly. It's like an alien code, spaghetti. But for some reason, it generalizes really well… It generalized better than our, we would argue, more elegant, hand-tuned harness.”

Machine Learning Street Talk · AI Research & Frontier Labs · September 2026

“We see the code it generated exactly. It's like an alien code, spaghetti. But for some reason, it generalizes really well… It generalized better than our, we would argue, more elegant, hand-tuned harness.” — Zhengyao Jiang, Machine Learning Street Talk

Jiang's team left a research agent running to improve its own harness, then tested the result on held-out benchmarks including weather forecasting — far outside anything it had been optimized for. The code it produced was unreadable and beat the harness his team had designed by hand. He admits they are torn between sticking with the clean, first-principles version and living with the one that works better.

Transcript

Machine Learning Street Talk Around 09:08 into the episode
Zhengyao Jiang

This project actually started about one and a half years ago, where we get interested in the concept of recursive self-improvement. I mean, historically, all of the research have diminishing returns. There is an idea, though, what if you can allow the researcher improve its own efficiency? Historically, it's possible by methodology improvement, better tools, but human researchers, like human researchers' brain, for example, is always the key bottleneck. If we now have an autonomous research system, can we point the research topic to itself so that it can improve its research efficiency? And the hope is one day this improvement of efficiency can overcome the diminishing return. So that, yeah, the curve of effort output turn from concave to convex. Of course, that's like the grand goal, like AGI. No one knows how far we can go, but improving the efficiency of self-referential research is definitely already on the roadmap.

Tim Scarfe

So let me set this up properly, right? So Weco, they created this agent called AID, and it is a research harness, right? And it's very, very simple. You give it a metric, you give it a problem, and it will hill climb towards that problem. But what they did was they set it up recursively. So they actually like hill climbed towards improving the agent itself. And what that means is improving the harness. So, you know, the prompts and the tooling and, you know, all of the skill surface, all of the things that the harness can do. And they just left it running. They just left it running and they saw what happened.

Zhengyao Jiang

There's an interesting thing we observe on self-evolved auto-research harness. We see the code it generated exactly. It's like an alien code, spaghetti. But for some reason, it generalizes really well. Of course, we have a set of benchmarks we're trying to hear Kaiber on. After that, we will test it on holdout benchmarks. Those are public benchmarks, like MLE Bench Light, ARE Bench Light. ARE Bench is like algorithmic discovery benchmark from Sarkana. MLE Bench is the OBI's machine learning engineering benchmark. And we also test it on a very far out of distribution task called Weather Bench 2. Essentially, you're trying to build a physics-based forecasting engine. To predict weather. And the weather bench is very far out of distribution from the training or optimization set of. But it still generalized really well. It generalized better than our, we would argue, more elegant, hand-tuned harness. We are a bit tall on the experiment results. Like to what extent we want to stick to more first principled and clean solutions we generated versus the solution that actually works better in practice. There's definitely space that we can improve over there, I think, like adding certain constraints to the optimizer. Two set of constraints might be something like you regularize neural networks. How do you make it trying to find the most simple solution to the same problem? That's one direction. But another direction might be, okay, you just need to live with it because one of the burdens that make people feeling it's a specatic code might be how much work the agent is done in the background. Like it do 100 experiments in our case on auto research self-improvement. And that number of experiments is probably similar amount that we did in the last two years. So if you onboard a new intern to look at our code base, maybe they also feel like, okay, this is so difficult to understand. I think part of it is, yeah, you need to bear with that cognition recognition burden of understanding huge amount of experiment results.

Keith Duggar

Do we even need those constraints at all?

Zhengyao Jiang

That's a really good question. I think in practice, you still need those constraints. Human and AI need to collaborate. There are certain aspects of human intelligence are still superior than AI. I believe so. For example, in Aiden, where we send the agent to attend OpenAI's hiring challenge, we found human is still much better at generating creative primitives. We found it's very important for human to build the first prototype that gives the agent the right search space. Basically, the initial abstraction will anchor the search space. A bit like how you design a neural network architecture, where you introduce inductive biases for the learning process. If the initial code base is based on a search-based scaffold, like aid, then a lot of ideas agent came up are actually coming from search literature. But if you give the initial code base as a React agent, React agent basically means you put all the context into the history and give the agent full flexibility, then a lot of the ideas come from the agent, the outer loop agent. Outer loop meaning the agent optimizing another agent.

Tim Scarfe

So what do we mean by recursive self-improvement, right? Are we talking about building a machine God? Are we just talking about some code that runs in a loop and optimizes itself? We try to pin this down.

Speaker names from our own diarization · position estimated from where the line sits in the episode

More from Machine Learning Street Talk