The Latency Hiding Burden

How SIMT and Domain Flow architectures handle memory latency differently

GPU: Register File Reads (N³)
0
Data Movement Ratio
KPU: L1 Buffer Reads (N²)
0
GPU (Stored Program)
Cycle: 0
MACs: 0
DONE
GEMM 1024³:
0.000%
Warp Scheduler Selecting warp...
Register File (64K) - SIMT Bottleneck
RF Read Intensity 0 reads/cycle
SIMT: Each thread fetches operands → N³ reads
Mem
Bus
SMEM
~10 cyc
90%
L2
~75 cyc
8%
DRAM
~200 cyc
2%
Active Warps
1/32
Waiting
28
Bank Conflicts
0
Efficiency
12%

The Problem: SIMT threads compute independent dot products. Each thread fetches its own A row and B column from the register file. For a 1024³ matmul: 3 billion RF accesses (N³ scaling). The register file becomes the energy bottleneck.

KPU (Domain Flow)
Cycle: 0
MACs: 0
DONE
GEMM 1024³:
0.000%
Credit-Based Pipeline Credits UP, Data DOWN
Host Memory
Credits: 2
L3 Buffer
Credits: 2
L2 Buffer
Credits: 2
L1 Stream
Credits: 2
A
B
C
↑ Credits flow back when buffer consumed ↑
16×16 Systolic Compute Tile (256 PEs)
L1 Buffer Activity 32 reads/cycle
Systolic: Operands enter once, reused N times → N² total reads
A operand reuse
16×
B operand reuse
16×
Pipeline Stages
4/4 Active
Stalls
0
Scheduler
None
Efficiency
94%

The Solution: Systolic schedules reuse operands as they flow through. Each A element is used by 16 PEs (one row). Each B element by 16 PEs (one column). For a 1024³ matmul: ~2 million L1 reads (N² scaling). That's 1500× less data movement than SIMT.

Speed: 1x
Ready (tile loaded)
Waiting for Memory
Registers Held
RF Reading
Data Flow
Credit Return