AI & Tech
GPUs Are Arrays of Tiny TPUs with Different Trade-offs
Dwarkesh Patel Podcast
Reiner Pope – Chip design from the bottom up
"At a very high-level point of view, the GPU has a lot of tiny, tiny TPUs tiled across the whole chip. You can sort of think of scaling this thing down into a really tiny unit with a smaller matrix unit, smaller vector unit. And that is sort of what an SM is."
Pope provided a novel architectural comparison showing that GPU streaming multiprocessors are essentially miniature TPUs. The key difference is that GPUs use many small units enabling higher bandwidth between vector and matrix operations through parallel wiring, while TPUs use fewer large units that amortize register file costs better but create data movement bottlenecks.
From this episode
Dwarkesh Patel Podcast