Four-bit floating-point (FP4) Tensor Cores promise to accelerate matrix multiplication, yet the overheads of scale computation, operand packing, and layout construction can often erase these gains. New research introduces "format-aware fusion," a co-design approach that optimizes quantization producers alongside their scale domains and consumer layouts to reclaim this lost efficiency.
Performance Benchmarks
The study evaluated the pretraining of Llama-3-family 8B models using 160 billion tokens. In matched accelerator probes, the performance gains were substantial:
- Standard bfloat16 reached 18.8K tokens/s/GPU.
- Transformer Engine NVFP4 achieved 27.6K tokens/s/GPU.
- The fastest custom format-aware fusion route reached 37.9K tokens/s/GPU.
Using MXFP4 with row-gradient stochastic rounding and fixed-sign 32-value Hadamard weight-gradient preconditioning, the system reached 37.2K tokens/s/GPU, representing 86.3% of bfloat16 model FLOP utilization.
Accuracy Trade-offs
While speed increased, training precision showed variations. The MXFP4 variant ended 2.11% above the raw bfloat16 training-loss endpoint. However, a hybrid approach using a Transformer Engine recipe with four final bfloat16 blocks narrowed this gap to 0.87% while maintaining a speed of 27.1K tokens/s/GPU.
The researchers noted that downstream rankings differ from training-loss rankings, indicating that FP4 outcomes are jointly dependent on the scale contract, the specific operand, and the execution path chosen during training.