Energy Breakdown by Component
I-Fetch 18%
Decode 10%
Sched 12%
Reg 30%
Op.Coll 8%
Compute 12%
Misc 10%
| Component | Energy |
| Instruction Fetch (I-Cache read) | 18 pJ |
| Instruction Decode | 10 pJ |
| Warp Scheduling | 12 pJ |
| Register File Access (3× per op) | 30 pJ |
| Operand Collection | 8 pJ |
| FMA Compute | 12 pJ |
| Misc (clock, etc.) | 10 pJ |
| Total per FMA | 100 pJ |
The GPU Tax: Every single FMA operation costs ~100 pJ,
but only 12 pJ does useful work. 88% of energy is wasted on
instruction infrastructure - fetching, decoding, scheduling,
and shuffling data through the register file.
Energy Breakdown by Component
Compute 80%
Data 12%
Cr 3%
5%
| Component | Energy |
| FMA Compute (systolic) | 12 pJ |
| Local Register (shift) | 0.5 pJ |
| Data Movement (amortized) | 1.5 pJ |
| Credit State Machine | 0.2 pJ |
| Misc (clock, etc.) | 0.8 pJ |
| Total per FMA | 15 pJ |
The Domain Flow Advantage: No instructions to fetch or decode.
No scheduler. No banked register file. Data flows directly through compute.
83% of energy does useful work. Same compute for 6.7× less total energy.