AI Research & Frontier Labs · August 2026

“The world is infinitely complex, and any simulation of it is like microscopic.” — Rich Sutton, Training Data

His answer to the synthetic-data strategy every frontier lab is betting on. Sutton's big world hypothesis says the world is massively more complex than any agent or any simulator — so an agent that cannot keep learning from its own experience is capped from the start.

Transcript

Training Data Around 14:28 into the episode
Sonya Huang

But you can have infinite synthetic worlds. The existing world is finite.

Khurram Javed

But let's go back to the echolocation thing, right? That's what I want. I want a drone that can localize itself and move with echolocation. If that's my goal, the robot, that's a robot, it's generating its own experience. So it could totally learn from its own experience, but it wouldn't be able to. It doesn't matter how much synthetic data you generate. It doesn't matter if you generate synthetic data that captures 50 different universes. It would not allow you to do that task without humans figuring it out first.

Rich Sutton

But first, just say it's the synthetic data is wrong. I mean, it won't be correct. It'll be a synthetic world. It won't be the real world. And it will matter. The world is incredibly complex. If you write a little program, because this is going to be a little program that will generate the synthetic data, it'll be a small world. So, for example, what's important to me is what's going on in your mind right now. Okay, and you're saying, why don't I get some synthetic data to tell me what's going on in other people's minds? No, there's no way we can have synthetic data for other people's minds. And other people's minds matter to us. You know, like I talked to you guys about investing today, so I care what you're going on. Your minds, and how could I get synthetic data on such a thing? Really, you can't even get synthetic data on anything. You can't get synthetic data on how the drone is going to interact with its environment, in the physical world, and the friction and wear in the motors of this robot. The world is infinitely complex, and any simulation of it is like microscopic. The big world hypothesis, let's say what it is, is that the world is massively more complex than your mind, than any agents, any agent. And this is obvious because the world contains many other agents. So, because the world is massively complex, there's no way you can do anything like anything that might claim to be optimal or perfect. You're going to be imperfect, and you have to have approximations, and those approximations will be severe. And so, because of that, that is ultimately the reason why we have to continue learning. If you want to think of it as a reason, we have to continue learning because we'll encounter some particular part of this immense world, and we'll have to learn an approximation that's tuned to the part of the world we're in, not to all the other parts that we're not in.

Sonya Huang

Yeah. I'm going to push on this one more time. And sorry, I'm being argumentative for the sake of being argumentative, but I'm trying to understand. My understanding is that the newest cohort of self-driving car companies, many of them were primarily trained in sim, and then they, you know, have to do some sort of post-training, I guess, to make sure they work in the real world, but that it's been a very effective pipeline.

Khurram Javed

Yeah. So, I think the important question to ask here is: how many engineers were involved in building that simulation? And are we ready to say that the only problem worth solving are those where we can hire a large team of engineers to first make a simulation? And I'm sure they had to do multiple iterations where they made the simulation, they learned in it, they realized there was a symptom gap that was not acceptable, then they fixed it. So, there is this human in the loop fixing the simulation. Like they're getting feedback from the real world, humans, and then they're fixing the simulation. Why can't we just remove the human and let the agent do it itself? And then, when

Rich Sutton

it actually drives, again, something unexpected will happen.

Speaker names from our own diarization · position estimated from where the line sits in the episode

More from Training Data