Dwarkesh Podcast · 22 May 2026 · 81 min

Reiner Pope – Chip design from the bottom up

AI chip designLogic gatesMultiply accumulateMatrix multiplicationFloating point arithmeticCircuit sizeData movementRegister filesSystolic arraysChip clock speedPipeline registersFPGA designASIC designCPU architectureGPU architectureTPU architectureCache memoryScratchpad memoryBranch predictionEnergy consumptionLow precision arithmetic

This episode delves into the fundamental workings of AI chips, beginning with basic logic gates and the multiply-accumulate operation, which is central to matrix multiplication. It explores the efficiency gains from architectural innovations like systolic arrays, contrasting them with traditional CPU/GPU designs and highlighting the critical role of optimizing compute-to-communication ratios. The discussion also covers chip synchronization via clock cycles, the trade-offs between FPGAs and ASICs, and the impact of memory architectures like caches versus scratchpads on performance and determinism.

Listen on Hopper →