Dwarkesh Podcast · AI Research & Frontier Labs · May 2026
Restating the GPU-versus-TPU architectural tradeoff in his own words, after Pope walked through how systolic arrays scale. The framing makes it a question about the workload rather than the hardware: irregular work rewards many small flexible units, uniform work rewards one big one.
So you're suggesting the tensor core within a streaming SM is analogous to an MXU. Yeah,
it's very, very similar. Yeah.
I see. And so if you had more lack of structure, having a bunch of tiny TPUs makes a lot of sense. Whereas if you kind of just have huge matrix multiplications, you're like, why don't we just avoid the cost of having the individual SMs with their own registers and work schedulers and things like that? Why don't we just make a huge thing and amortize those costs across the whole thing?
And I mean, I think this shows up in how large you can grow things. We've sort of seen this theme, especially with a systolic array, where larger systolic array amortizes the register file costs better. This sort of design allows you to have larger systolic arrays, whereas the sort of GPU design constrains you to having small units of everything. There is a trade-off, however. There ends up being, because of this sort of coarse-grained separation of things, you need to move a lot of data from the vector unit to the matrix units. And so you need to move a lot of data through a sort of like two lines of parameter here. Whereas if you sort of look at the equivalent thing here, you've got vector units everywhere, and you need to move data through this line, through this line, through this line, through this line, through this line, through this line. So the amount of data you can move between a vector unit and a matrix unit is actually much higher in a GPU than in a TPU because instead of having to move all the data through these just two lines, you're moving all these data through 16 lines or something of wiring instead in a GPU.
Right, but also you might have to move across less area.
Which I mean is also saving its energy sensor. So data ends up moving like if you can operate entirely within an SM, the data movement is much smaller. But then the moment you want to operate across SMs, it becomes sort of more complicated and expensive. So