DP Dwarkesh Patel Dwarkesh Patel Hosts the Dwarkesh Podcast, a long-form interview show on AI progress, timelines, and the people building frontier models, along with economists and historians.

“It's interesting to me that when we were talking last time about inference across many chips, the big high-level thing we're trying to optimize for is increasing compute per memory bandwidth — that is to say, per communication. And here also we're trying to increase the amount of actual multiplies or additions relative to transporting information from registers to the logic. So in both cases you're trying to maximize compute relative to communication.”

Dwarkesh Podcast · AI Research & Frontier Labs · May 2026

“It's interesting to me that when we were talking last time about inference across many chips, the big high-level thing we're trying to optimize for is increasing compute per memory bandwidth — that is to say, per communication. And here also we're trying to increase the amount of actual multiplies or additions relative to transporting information from registers to the logic. So in both cases you're trying to maximize compute relative to communication.” — Dwarkesh Patel, Dwarkesh Podcast

Reiner Pope was explaining why you'd load a value into a systolic array slowly over narrow lanes, since bandwidth costs die area. Dwarkesh's move is to notice the same tradeoff he'd learned one episode earlier at a completely different level of the stack — across chips rather than across gates. Pope confirms it shows up all the way up and down.

Transcript

Dwarkesh Podcast Around 28:57 into the episode
Dwarkesh Patel

Would you mind repeating this sentence,

Reiner Pope

so like we're sort of like we know that we're going to be bringing in numbers only rarely into the matrix. And so we just want to come up with any construction at all such that the amount of wiring that actually feeds into sort of crosses this boundary of the systolic array, like this boundary right here, we just want to keep that bounded to x and not go as xy. And so a particularly simple strategy is that we sort of bring in a number into the top row of the systolic array. That's what we can do in one clock cycle. And then for y consecutive clock cycles, we're going to be bringing in the top row every time and then sort of shift all of the other rows down by one. And that keeps the wiring that needs to come from this expensive register file only down to a factor of x rather than xy.

Speaker 3

I see. Okay. So there's two questions in terms of communication. There's like communication time and then there's communication bandwidth. Yes. And you're saying, since we're only going to be loading this in once, let's minimize bandwidth. That's right. Because bandwidth equals die area. And let's just load it in slowly over like smaller lanes because we're just going to keep this value in there for a while. Exactly. Interesting. So it's interesting to me that when we were talking last time about inference across many chips, the big high-level thing we're trying to optimize for is increase the amount of compute per memory bandwidth, that is to say, per communication. And here also, we're trying to increase the amount of like actual multiplies or actual additions relative to transporting information from registers to the logic. So in both cases, you're trying to maximize compute relative to communication.

Reiner Pope

Yeah. This shows up sort of all the way up and down the stack. This is sort of close to the bottom, sort of like to the gates. There's sort of a version that's maybe even closer to the gates of just like even the precision of number format that you choose to use. We saw that same effect. There's like a square cube law or a like squared versus linear term going on, both in just purely the precision of this ALU, but then also in terms of the size of the matrix.

Speaker 3

Yeah. Very interesting.

Reiner Pope

So this unit is sort of the next bigger unit. We had like the multiplication circuit. And then on top of that, we have a pretty large systolic array. I drew it as two by two, but in like, for example, older TPUs, they were described as 128 by 128 of this circuit shown here. And this circuit ends up being, this is the most efficient known mechanism for a circuit for implementing a matrix multiply.

Speaker names from our own diarization · position estimated from where the line sits in the episode

More from Dwarkesh Podcast