Register File
64K Regs
Warp
Scheduler
Operand
Collector
I-Cache
Decode
Tensor
Cores
Shared
Memory
L2 Cache
Memory Controllers (HBM/GDDR)
Thread Block Scheduler & Dispatch
Clock, Power, Misc
Click a region to learn more
Most silicon is dedicated to latency hiding infrastructure: register files, schedulers, and caches.
Area Breakdown
28%
15%
8%
10%
7%
12%
12%
8%
The GPU Dilemma: To hide memory latency, GPUs need massive register files
(64K+ registers), complex schedulers, and deep instruction caches.
Only ~12% of silicon does actual computation.
Systolic Compute Tiles
domain flow array
L3
Buffers
L2
Buffers
L1
Credit
Control
DMA Engine
NoC
Memory Controllers (LPDDR5/HBM)
Click a region to learn more
Most silicon is pure compute. Small buffers replace massive register files. No scheduler needed.
The Domain Flow Advantage: System-level scheduling eliminates the need for
register files and schedulers. 72% of silicon is pure compute.
Smaller die = lower cost, lower power, higher yield.