Energy Flow: Where Your Power Goes

Understanding energy efficiency in SIMT vs Domain Flow architectures

GPU (Stored Program)
12% Efficient
INPUT
100 pJ
Useful
Compute
12 pJ
Overhead
88 pJ
Energy Breakdown by Component
I-Fetch 18%
Decode 10%
Sched 12%
Reg 30%
Op.Coll 8%
Compute 12%
Misc 10%
Instruction Fetch
Decode
Scheduling
Register Access
Operand Collect
Compute
ComponentEnergy
Instruction Fetch (I-Cache read)18 pJ
Instruction Decode10 pJ
Warp Scheduling12 pJ
Register File Access (3× per op)30 pJ
Operand Collection8 pJ
FMA Compute12 pJ
Misc (clock, etc.)10 pJ
Total per FMA100 pJ

The GPU Tax: Every single FMA operation costs ~100 pJ, but only 12 pJ does useful work. 88% of energy is wasted on instruction infrastructure - fetching, decoding, scheduling, and shuffling data through the register file.

KPU (Domain Flow)
83% Efficient
INPUT
15 pJ
Useful
Compute
12 pJ
Overhead
3 pJ
Energy Breakdown by Component
Compute 80%
Data 12%
Cr 3%
5%
Systolic Compute
Data Movement
Credit Logic
Misc
ComponentEnergy
FMA Compute (systolic)12 pJ
Local Register (shift)0.5 pJ
Data Movement (amortized)1.5 pJ
Credit State Machine0.2 pJ
Misc (clock, etc.)0.8 pJ
Total per FMA15 pJ

The Domain Flow Advantage: No instructions to fetch or decode. No scheduler. No banked register file. Data flows directly through compute. 83% of energy does useful work. Same compute for 6.7× less total energy.

Energy per FMA Operation
100pJ
15pJ
6.7× less energy
Energy Efficiency
12%
83%
6.9× more efficient
Ops per Watt (Relative)
1×
6.7×
Same compute, fraction of power