DP Dwarkesh Patel Dwarkesh Patel Hosts the Dwarkesh Podcast, a long-form interview show on AI progress, timelines, and the people building frontier models, along with economists and historians.

“When I had Dario on the podcast, the thing I asked him was: if you truly expect models which will be human-like in their ability to learn on the job, why would you try to bake in all these skills of working with PowerPoint or something? Why not just expect the model to be able to pick that up while it's deployed?”

Dwarkesh Podcast · AI Research & Frontier Labs · September 2026

“When I had Dario on the podcast, the thing I asked him was: if you truly expect models which will be human-like in their ability to learn on the job, why would you try to bake in all these skills of working with PowerPoint or something? Why not just expect the model to be able to pick that up while it's deployed?” — Dwarkesh Patel, Dwarkesh Podcast

Pressing the panel on why labs still invest so heavily in domain-specific RL training when their own stated bet is on models that learn continually. Schulman's answer is that if models were good enough at in-context learning you wouldn't need to train them on finance — which is the point, and also the concession.

Transcript

Dwarkesh Podcast
Dwarkesh Patel

yeah yeah maybe taking a step back here's what i here's what it seems to me that the plan uh for ai research going forward is and you tell me if you think it's going to work or if you agree with this characterization so the bet is that we will scale up our all VR training across millions of diverse environments across hundreds of different kinds of domains and what will emerge at the other end is an agent which has like learned these basic skills or less than basic skills around being persistent being able to triage information and context eventually having like end-to-end optimization of working with other agents and things like that and such an agent will be very sample efficient within the context you know you've done research on how you actually scale up in context learning to make it like arbitrarily long but you keep scaling it up and so what comes out the other end will something will be something that it basically functions like a drop in remote worker over the course of a week or a month first of all do you agree that that is the bet the labs are making and second is it is that enough like basically learning how to learn within the simulacra within a data center uh and then but getting deployed into the real world But not actually learning from real-world deployment, only learning these meta-skills from the simulated environments in the data center.

Charlie O'Neill

Yeah, I think it's now hard to separate out how much of the lab's effort is going towards direct RSI versus making generally intelligent models they can continue to deploy to collect revenue to fund the next big training room. I think for the latter, yes, that's probably just the bet they're making. And it's very clear, like the pattern of where these environments are going over the last few years. I mean, like Anthropic's lineage of environments is a very clear example of this. First, we just focus on coding and we're going to get really, really good at that. And then the task horizon that we've got from coding, which is probably the lowest hanging fruit in terms of data available on the internet to create environments, like their own internal stuff that they can turn into environments. Then we're going to generalize, we're going to go after finance next. And literally just so much Excel data and all that sort of stuff in the RL training. And then it's PowerPoints. It's this long tail of the working economy. And that seemed to work really well. And a lot of the other labs, I think, even the open source labs have now realized that that was the correct answer. But

Dwarkesh Patel

what is the implication from that? When I had Dario on the podcast, the thing I asked him was: if you truly expect models which will be human-like in their ability to learn on the job, why would you try to bake in all these skills of working with PowerPoint or something? What did you just expect the model to be able to pick that up on while it's deployed? And so, yeah, there's multiple different explanations. One is just that this is, we expect models to get there soon, but they're not there yet. So why not amortize these skills into the model training? Another is that we're not concentrated on making it really good at widely deployed work. We just want it really good at RSI. And this is just a way for us to get revenue so that we can pour it back into a model that is actually really good at doing RSI development. And then once the singularity happens, the thing that comes out the other end will be really good at all the things which seem like bottlenecks to the current generation of models. Yeah, John, I don't know if you have takes on how once you construe why there is so much task-specific knowledge in these models, if the path is this kind of generalization.

John Schulman

Yeah. I mean, if the models were good enough at learning in context, then in theory, you wouldn't need to train them on finance. They would just be able to figure out, read all the books on the fly and figure out how to do everything in the appropriate jurisdiction. Yeah, and you could argue that you need to do a lot of this domain-specific training just to make them more efficient. So even if they were smart enough to figure this out on the fly, you still might want to do a bunch of RL and bake all these intuitions into the weights so the model would be more efficient at runtime. Yeah, I'd say in practice, it does seem like model providers are going domain by domain and trying to strengthen the models in the highest value domain. And I'd say that that's one of the answers to why the models have gotten so much better. It's just because the model providers have covered a lot of the high-value domains and the most common types of skills. I mean,

Beren Millidge

I think another thing is just that it's not that expensive to do both at the same time, right? Because the models are massive. They can easily afford in terms of their parameters to learn everything. And there is likely some transfer and sort of even just even if finance is not specifically like the information is important for like RSI, just the general meta-learning of how to figure out what's important, how to have taste, how to do long horizon work is potentially generalizable. And there's not that much RSI like data in the world as well. It's kind of hard to generate and that requires a lot of effort. So if you can sort of amortize in this other data, get some transfer from it, you already have massive compute and massive parameter space. So why not do that as well as obviously the direct commercial intent of selling your model?

Dwarkesh Patel

That makes sense. Oh,

Speaker names from our own diarization · position estimated from where the line sits in the episode

More from Dwarkesh Podcast