AI Research & Frontier Labs · September 2026
The panel had been arguing about whether reinforcement learning flattens the range of a model's writing, with Schulman noting that RL-trained models reuse the same themes and character names. He extends that from one model to the whole open-weight ecosystem. Beren Millidge pushes back immediately, arguing this is a property of the training data rather than of RL or distillation as methods.
analysis, you find that they're reusing certain themes all the time. And they're using the same character names all the time. So there's actually, it's not like you're getting the same kind of diversity that you get when you, like from human authors, you're sort of getting one really good style. So I think that kind of
diversity has
definitely been cut down by RL a lot. And in fact, since we were talking about distillation earlier, that's sort of something, yeah, one thing that's happening is that so many people are distilling mostly from Claude that all the open weight models write the same way as Claude and use the same, like have the same ticks. So this seems kind of concerning to me that we're having this monoculture emerge. Yeah.
Again, I don't think this is fundamental to RL as a method though. And same with distillation. Even with distillation, you're just training on the data. Just because your data is not super broad, that doesn't mean the training method itself is somehow wrong. It's a problem with the data. And I think a lot of, for instance, the RL entropy collapse is basically due to exploitation of fairly simple verifiers when you don't have a huge diversity of environments. Because for instance, the writing, I think the writing is presumably graded by some judge and the judge has some specific ticks and the model is learning to award hack the judge and that's why like it collapses. But this is really a problem with the judge, it's not a problem with RL in general.
Okay, super rapid fire predictions about the future. So I want timelines on the following couple questions. By when do we have models which you can here's what the it feels like to a user you basically hire them as a drop-in remote worker for all kinds of white-collar work, not just coding, but I don't know, video editing, law, paralegal, et cetera. It's like literally an actual remote worker with full computer use with literally a month of seamless learning and operation and executing on like complex projects and it required interacting with other people, et cetera, et cetera. It's like everything a human worker could do over a month.
If you mandate it to use a browser or whatever, rather than, again, the firm setting up the information to be programmatically accessible, maybe a couple of years. But if it's not browser-based, it can send Slack messages, it can do all this stuff. But I'd still probably say it around a year.