RS

Rich Sutton

Things Rich Says on Podcasts

Argues that agents unable to keep learning from experience after deployment are capped from the start, a view described as the big world hypothesis: the world is too complex for any simulation to fully capture, so continual learning matters more than better pretraining. Also known for the bitter lesson, the observation that general methods that lean on computation tend to outperform approaches built on hand-designed features. Professor of computing science at the University of Alberta and chief scientific advisor at the Alberta Machine Intelligence Institute. Co-author of the textbook Reinforcement Learning: An Introduction and co-recipient, with Andrew Barto, of the 2024 ACM A.M. Turing Award for foundational work in reinforcement learning. Appeared on the Training Data podcast.

Where to Find Them

Rich Sutton writes Eye On A.I. . They have also been a guest on Training Data .

Recently: “Rich Sutton and Khurram Javed: Why AI Models Stop Learning, and How to Start It Again” on Training Data (August 2026); “Rich Sutton Edit V5-Norm 01-01” on Eye On A.I. (March 2019).

What They Said

“The world is infinitely complex, and any simulation of it is like microscopic.” — Rich Sutton, Training Data

His answer to the synthetic-data strategy every frontier lab is betting on. Sutton's big world hypothesis says the world is massively more complex than any agent or any simulator — so an agent that cannot keep learning from its own experience is capped from the start.

Training Data · 2026-08-18 Permalink → Listen →
Training Data Around 14:28 into the episode
Sonya Huang

But you can have infinite synthetic worlds. The existing world is finite.

Khurram Javed

But let's go back to the echolocation thing, right? That's what I want. I want a drone that can localize itself and move with echolocation. If that's my goal, the robot, that's a robot, it's generating its own experience. So it could totally learn from its own experience, but it wouldn't be able to. It doesn't matter how much synthetic data you generate. It doesn't matter if you generate synthetic data that captures 50 different universes. It would not allow you to do that task without humans figuring it out first.

Rich Sutton

But first, just say it's the synthetic data is wrong. I mean, it won't be correct. It'll be a synthetic world. It won't be the real world. And it will matter. The world is incredibly complex. If you write a little program, because this is going to be a little program that will generate the synthetic data, it'll be a small world. So, for example, what's important to me is what's going on in your mind right now. Okay, and you're saying, why don't I get some synthetic data to tell me what's going on in other people's minds? No, there's no way we can have synthetic data for other people's minds. And other people's minds matter to us. You know, like I talked to you guys about investing today, so I care what you're going on. Your minds, and how could I get synthetic data on such a thing? Really, you can't even get synthetic data on anything. You can't get synthetic data on how the drone is going to interact with its environment, in the physical world, and the friction and wear in the motors of this robot. The world is infinitely complex, and any simulation of it is like microscopic. The big world hypothesis, let's say what it is, is that the world is massively more complex than your mind, than any agents, any agent. And this is obvious because the world contains many other agents. So, because the world is massively complex, there's no way you can do anything like anything that might claim to be optimal or perfect. You're going to be imperfect, and you have to have approximations, and those approximations will be severe. And so, because of that, that is ultimately the reason why we have to continue learning. If you want to think of it as a reason, we have to continue learning because we'll encounter some particular part of this immense world, and we'll have to learn an approximation that's tuned to the part of the world we're in, not to all the other parts that we're not in.

Sonya Huang

Yeah. I'm going to push on this one more time. And sorry, I'm being argumentative for the sake of being argumentative, but I'm trying to understand. My understanding is that the newest cohort of self-driving car companies, many of them were primarily trained in sim, and then they, you know, have to do some sort of post-training, I guess, to make sure they work in the real world, but that it's been a very effective pipeline.

Khurram Javed

Yeah. So, I think the important question to ask here is: how many engineers were involved in building that simulation? And are we ready to say that the only problem worth solving are those where we can hire a large team of engineers to first make a simulation? And I'm sure they had to do multiple iterations where they made the simulation, they learned in it, they realized there was a symptom gap that was not acceptable, then they fixed it. So, there is this human in the loop fixing the simulation. Like they're getting feedback from the real world, humans, and then they're fixing the simulation. Why can't we just remove the human and let the agent do it itself? And then, when

Rich Sutton

it actually drives, again, something unexpected will happen.

Speaker names from our own diarization · position estimated from where the line sits in the episode
“I'm not weird. The field is weird. The field, they need to call it continual learning. It's just learning.” — Rich Sutton, Training Data

He opens the episode with it and returns to it word for word an hour later. People keep telling Sutton his position is radical; he thinks the rest of the field started thinking strangely about a decade ago — before the AI boom, nobody needed the phrase "continual learning" because no other kind existed.

Training Data · 2026-08-18 Permalink → Listen →
Training Data Around 00:00 into the episode
Rich Sutton

People think I have a radical point of view sometimes. They start questions saying how what I'm thinking is so different from everyone else. But I don't see it that way at all. I see it as like I'm thinking the ordinary way. It's just everyone else that's thinking a bit weird. And I mean that, like, you know, it's just the recent times people are thinking weird. Before there was all this AI craziness, you talk about, you wouldn't have to say continual learning because it wouldn't make any sense to talk about learning that wasn't continual. All learning is continual. We always act and we learn. That's just the normal way of thinking. I'm not weird. The field is weird. The field, they need to call it continual learning. It's just learning.

Sonya Huang

We are honored to have the great Rich Sutton with us here today. Rich, you invented reinforcement learning. You wrote the seminal textbook. You're the key students in the field, folks like Dave Silver. You wrote the essay, The Bitter Lesson, that I believe is the Bible of the field. And you have just been one of the greats in propelling the field forward. So thank you for taking the time to join us today. Rich is joined by Kuram Javed, his co-founder and former student from the University of Alberta. The two of you have set off to found Oak Lab. I'm very excited to talk to you about that today. So for today's session, we're going to start talking about the Bitter Lesson, the state of the world as we know it today, whether LLMs will get us there or not. And then we're going to transition to start talking about your research agenda and your plan for Oak. Rich, maybe take us back. I was going to start with the Bitter Lesson, but I actually want to start earlier than that. Decades ago, you decided to dedicate your career to reinforcement learning, to deep reinforcement learning in particular, and you established the University of Alberta as a bastion of that back when I think the field was very much in its infancy. What gave you the conviction to do that?

Rich Sutton

What else are you going to

Sonya Huang

do?

Speaker names from our own diarization · position estimated from where the line sits in the episode

Collections They Appear In