DP Dwarkesh Patel Dwarkesh Patel Hosts the Dwarkesh Podcast, a long-form interview show on AI progress, timelines, and the people building frontier models, along with economists and historians.

“If you had more lack of structure, having a bunch of tiny TPUs makes a lot of sense. Whereas if you just have huge matrix multiplications, you'd think: why don't we just avoid the cost of individual SMs with their own registers and work schedulers, and make one huge thing that amortizes those costs across the whole thing?”

Dwarkesh Podcast · AI Research & Frontier Labs · May 2026

“If you had more lack of structure, having a bunch of tiny TPUs makes a lot of sense. Whereas if you just have huge matrix multiplications, you'd think: why don't we just avoid the cost of individual SMs with their own registers and work schedulers, and make one huge thing that amortizes those costs across the whole thing?” — Dwarkesh Patel, Dwarkesh Podcast

Restating the GPU-versus-TPU architectural tradeoff in his own words, after Pope walked through how systolic arrays scale. The framing makes it a question about the workload rather than the hardware: irregular work rewards many small flexible units, uniform work rewards one big one.

Transcript

Dwarkesh Podcast Around 1:16:53 into the episode
Dwarkesh Patel

So you're suggesting the tensor core within a streaming SM is analogous to an MXU. Yeah,

Reiner Pope

it's very, very similar. Yeah.

Dwarkesh Patel

I see. And so if you had more lack of structure, having a bunch of tiny TPUs makes a lot of sense. Whereas if you kind of just have huge matrix multiplications, you're like, why don't we just avoid the cost of having the individual SMs with their own registers and work schedulers and things like that? Why don't we just make a huge thing and amortize those costs across the whole thing?

Reiner Pope

And I mean, I think this shows up in how large you can grow things. We've sort of seen this theme, especially with a systolic array, where larger systolic array amortizes the register file costs better. This sort of design allows you to have larger systolic arrays, whereas the sort of GPU design constrains you to having small units of everything. There is a trade-off, however. There ends up being, because of this sort of coarse-grained separation of things, you need to move a lot of data from the vector unit to the matrix units. And so you need to move a lot of data through a sort of like two lines of parameter here. Whereas if you sort of look at the equivalent thing here, you've got vector units everywhere, and you need to move data through this line, through this line, through this line, through this line, through this line, through this line. So the amount of data you can move between a vector unit and a matrix unit is actually much higher in a GPU than in a TPU because instead of having to move all the data through these just two lines, you're moving all these data through 16 lines or something of wiring instead in a GPU.

Dwarkesh Patel

Right, but also you might have to move across less area.

Reiner Pope

Which I mean is also saving its energy sensor. So data ends up moving like if you can operate entirely within an SM, the data movement is much smaller. But then the moment you want to operate across SMs, it becomes sort of more complicated and expensive. So

Speaker names from our own diarization · position estimated from where the line sits in the episode

More from Dwarkesh Podcast