How SIMT and Domain Flow architectures handle memory latency differently
GPU: Register File Reads (N³)
0
Data Movement Ratio
1×
KPU: L1 Buffer Reads (N²)
0
GPU (Stored Program)
Cycle: 0
MACs: 0
DONE
GEMM 1024³:
0.000%
Warp SchedulerSelecting warp...
Register File (64K) - SIMT Bottleneck
RF Read Intensity0 reads/cycle
SIMT: Each thread fetches operands → N³ reads
Mem Bus
SMEM
~10 cyc
90%
L2
~75 cyc
8%
DRAM
~200 cyc
2%
Active Warps
1/32
Waiting
28
Bank Conflicts
0
Efficiency
12%
The Problem: SIMT threads compute independent dot products.
Each thread fetches its own A row and B column from the register file.
For a 1024³ matmul: 3 billion RF accesses (N³ scaling).
The register file becomes the energy bottleneck.
KPU (Domain Flow)
Cycle: 0
MACs: 0
DONE
GEMM 1024³:
0.000%
Credit-Based PipelineCredits UP, Data DOWN
Host Memory
Credits: 2
→
L3 Buffer
Credits: 2
→
L2 Buffer
Credits: 2
→
L1 Stream
Credits: 2
A
B
C
↑ Credits flow back when buffer consumed ↑
✓
✓
✓
16×16 Systolic Compute Tile (256 PEs)
L1 Buffer Activity32 reads/cycle
Systolic: Operands enter once, reused N times → N² total reads
A operand reuse
16×
B operand reuse
16×
Pipeline Stages
4/4 Active
Stalls
0
Scheduler
None
Efficiency
94%
The Solution: Systolic schedules reuse operands as they flow through.
Each A element is used by 16 PEs (one row). Each B element by 16 PEs (one column).
For a 1024³ matmul: ~2 million L1 reads (N² scaling).
That's 1500× less data movement than SIMT.