Dwarkesh Podcast · AI Research & Frontier Labs · May 2026
Reiner Pope was explaining why you'd load a value into a systolic array slowly over narrow lanes, since bandwidth costs die area. Dwarkesh's move is to notice the same tradeoff he'd learned one episode earlier at a completely different level of the stack — across chips rather than across gates. Pope confirms it shows up all the way up and down.
Would you mind repeating this sentence,
so like we're sort of like we know that we're going to be bringing in numbers only rarely into the matrix. And so we just want to come up with any construction at all such that the amount of wiring that actually feeds into sort of crosses this boundary of the systolic array, like this boundary right here, we just want to keep that bounded to x and not go as xy. And so a particularly simple strategy is that we sort of bring in a number into the top row of the systolic array. That's what we can do in one clock cycle. And then for y consecutive clock cycles, we're going to be bringing in the top row every time and then sort of shift all of the other rows down by one. And that keeps the wiring that needs to come from this expensive register file only down to a factor of x rather than xy.
I see. Okay. So there's two questions in terms of communication. There's like communication time and then there's communication bandwidth. Yes. And you're saying, since we're only going to be loading this in once, let's minimize bandwidth. That's right. Because bandwidth equals die area. And let's just load it in slowly over like smaller lanes because we're just going to keep this value in there for a while. Exactly. Interesting. So it's interesting to me that when we were talking last time about inference across many chips, the big high-level thing we're trying to optimize for is increase the amount of compute per memory bandwidth, that is to say, per communication. And here also, we're trying to increase the amount of like actual multiplies or actual additions relative to transporting information from registers to the logic. So in both cases, you're trying to maximize compute relative to communication.
Yeah. This shows up sort of all the way up and down the stack. This is sort of close to the bottom, sort of like to the gates. There's sort of a version that's maybe even closer to the gates of just like even the precision of number format that you choose to use. We saw that same effect. There's like a square cube law or a like squared versus linear term going on, both in just purely the precision of this ALU, but then also in terms of the size of the matrix.
Yeah. Very interesting.
So this unit is sort of the next bigger unit. We had like the multiplication circuit. And then on top of that, we have a pretty large systolic array. I drew it as two by two, but in like, for example, older TPUs, they were described as 128 by 128 of this circuit shown here. And this circuit ends up being, this is the most efficient known mechanism for a circuit for implementing a matrix multiply.