Dwarkesh Podcast · AI Research & Frontier Labs · May 2026
Teaching out his own 'bits per flop' framework, which decomposes into samples per flop times bits per sample. His point is that long-horizon RL is punished on both terms at once — fewer samples per unit of compute, and less information per sample than supervised learning's cross-entropy signal would give you.
Why have they converged to that?
It's just more stable. So you might use the off-policy Q as a way to do advantage computation, like Q minus sum of Q. That's kind of like your, or sorry, like, you know, sum of like if there's n actions, and then, yeah. So like this is your value, and then this is your kind of current Q values. Your advantage for that action is like the average value minus your current one. So like people can try to estimate Q in an off-policy way and then like just use advantage here. And then the sort of, if there's a problem in these dynamics, it doesn't blow up your loss as much. And so in robotics, there's a kind of convergence towards more like using off-policy data to just shape your rewards, but not actually be directly your.
I'm reminded now of our earlier conversation of why MCTS is so favorable as compared to the kind of reinforce or policy gradient kind of thing LLMs do. And this might be totally wrong, but I wrote a blog post a few months ago about how RL, at least policy gradient RL, is even more inefficient than you might think. And so the inefficiency one thinks about naively is the fact that you have to roll out a whole trajectory in order to get any learning signal at all. And so as these trajectories become longer and longer, as an agent has to, instead of just previously complete the next word in the sentence, it has to go instead to, hey, do two days worth of work to figure out even if you even did this project correctly, the amount of information per flop has been decreasing. As you had to unroll two days worth of thinking in order to see if you even did something correctly, to like, did I implement this feature? The amount of samples per flop has been decreasing. But so you can think of you're trying to maximize as your learning bits per flop, right? And this is, you can think of bits per flop as samples per flop times bits per sample. And what I just mentioned a second ago is that the samples per flop go down as RL becomes more and more along horizon. But at least this kind of naive RL is also terrible from a bits per sample perspective. And here's what I mean, at least compared to supervised learning. So early on in training, let's say you have a vocabulary for an LLM that is 100k long. So there's 100K possible tokens that one could answer. And you have a totally untrained model, and you have a prompt like the sky is with supervised learning, what would happen is that the model would have some probability distribution over all the things it could say. There's a label that says actually the term here is blue. And it would learn basically for cross-entributal loss exactly how far its distribution is from correctly saying blue. Now, if you're doing this through RL, you would say the model would try the sky is halicon. Nope, that's wrong. The sky is told. Nope, that's wrong. This is a totally untrained model, right? And so you would have to do this on the order of 100,000 times in order to just stumble on blue, then get some learning signal off of that. So if you're in the supervised learning regime and you just get, you have your distribution of probabilities, you get told that it's blue and you figure out how far off you are, the amount you learn is a function of your pass rate. So the further away you are from blue, the more you've learned to go towards blue. Using cross-entropy loss. And so you can think of it as your pass rate, your prior probability of having said blue. And as a function of that, in supervised learning, through cross-entropy loss, you would learn negative log p, p being pass rate bits once you get this label. Whereas in RL, if you're just randomly guessing shit and seeing if it works or not, that's just basically going to be the entropy of a binary random variable, which is
And what's also tough here is that actually the distribution that you're sampling under is your policies distribution. So it's like if your policy has no chance of sampling blue, then you will never get a signal.
Exactly, right, right. So that's being modeled by the fact that your probability of sampling blue is extremely low. If you do sample it, you do learn as much as you would have learned in a supervised learning. In all other cases, like 99.999% of in an untrained model, you're just learning incredibly little from like seeing halicon is not the correct word or told is not the correct word. And that's what happens most of the time. So you're just like, learn very little. So if you try to graph, if you put on the x-axis your pass rate, and here you put the sort of the bits, bits you're learning from a sample. If you have like 0% here, 50% here, and 100% here, so the end of trading, you're here. If you have supervised learning, negative log pass rate would look something like this. And then the entropy binary random variable would look like this. And this is, depending on whether you're doing NATs or bits. Yeah, if you do bits, it's like one, right? Here at the peak. This is like a coin flip. You learn the most from a coin flip. This is supervised learning. This is RL. However, the problem is you spend most of training in this regime, right? Like in the low pass rate regime. And in fact, if how fast you're learning is a function of how many bits per sample you're getting, and you're getting very little signal here. If you chart the pass rate on a log scale, so you put the x-axis on a log scale, where like at the beginning of training with a vocab size of 100K, the pass rate is 1 over 100,000, then 1 over 10,000, 1 over 1,000, 1 over 100, and then, okay, what this graph looks like here, where supervised learning would look like this, and then RL, if you just basically crunch what I just showed there, it would look like that.
Yeah, and arguably you spend all your time here, potentially never even getting a single success, right? Exactly. So it's a sort of depressing plot in the sense that once you're here, it's not at all obvious how you get to here. Once you're here, you have something, but you actually, in many RL problems, spend all the time here. So there's a sort of question of how do you initialize so you're at least not at zero, but at a non-zero pass rate. One more thing I'd like to add about bits per sample that's very relevant to any kind of machine learning problem is that there's a connection to soft targets and distillation where if you have access to the logits, right, not just the one hot, like this, this is a sort of one hot token answer. If you have access to the soft targets, the entropy of this distribution is far, far higher than the one hot. So there's actually way more, there's way more information and bits per sample in a soft label. So that's why distillation is so effective per sample, is that it's actually giving you way more information per sample.