Silicon Budget: Where the Transistors Go

Comparing chip area allocation between SIMT and Domain Flow architectures

GPU Die Layout
~400 mm² typical
Register File
64K Regs
Warp
Scheduler
Operand
Collector
I-Cache
Decode
Tensor
Cores
Shared
Memory
L2 Cache
Memory Controllers (HBM/GDDR)
Thread Block Scheduler & Dispatch
Clock, Power, Misc
Click a region to learn more
Compute Area
~12%
Overhead Area
~88%
Most silicon is dedicated to latency hiding infrastructure: register files, schedulers, and caches.
Area Breakdown
28%
15%
8%
10%
7%
12%
12%
8%
Registers
Schedulers
Operand Coll.
I-Cache
Tensor Cores

The GPU Dilemma: To hide memory latency, GPUs need massive register files (64K+ registers), complex schedulers, and deep instruction caches. Only ~12% of silicon does actual computation.

KPU Die Layout
~150 mm² equivalent
Systolic Compute Tiles
domain flow array
L3
Buffers
L2
Buffers
L1
Credit
Control
DMA Engine
NoC
Memory Controllers (LPDDR5/HBM)
Click a region to learn more
Compute Area
~72%
Buffer Area
~20%
Most silicon is pure compute. Small buffers replace massive register files. No scheduler needed.
Area Breakdown
72%
12%
6%
2%
8%
Compute Tile
Stream Buffers
DMA
Credit Logic
Memory Ctrl

The Domain Flow Advantage: System-level scheduling eliminates the need for register files and schedulers. 72% of silicon is pure compute. Smaller die = lower cost, lower power, higher yield.

Compute Density
12%
72%
6× more compute per mm²
Register/Buffer Area
28%
12%
2.3× less storage overhead
Scheduler/Control
23%
2%
11× less control logic
Die Size (Equivalent Compute)
400mm²
150mm²
2.7× smaller die