Researchers Closed Gap in Native 8-Bit LLM Training
A new method restores mathematical precision in FP8 training, enabling compute gains without accuracy loss.
Updated on Oct. 1, 2026 in Artificial Intelligence

Live Poll
Do you believe advancements in AI efficiency will significantly lower the cost of building large models?
Researchers from MIT and Carnegie Mellon University have identified the root cause of accuracy degradation in 8-bit large language model training. Their new method, Delta-Matching, achieves parity with full-precision baselines during model development.
Why it matters
The development potentially unlocks the full 3,958-teraflop performance of NVIDIA H100 hardware, doubling the compute efficiency previously available in FP8. By solving systematic gradient bias, it enables more efficient training without requiring architectural changes.
Delta-Matching scales 8-bit performance to match BF16/FP32 baselines, overcoming the 16.3% benchmark scores observed in naive implementations. This approach fully utilizes the 3,958-teraflop capacity of NVIDIA H100 GPUs.
The players
MIT
A research university known for advanced computer science, artificial intelligence, and hardware architecture.
Carnegie Mellon University
A global leader in machine learning, robotics, and artificial intelligence research.
NVIDIA
A semiconductor company providing the H100 GPU architecture, which dominates the market for large-scale AI training.
The details
Delta-Matching addresses a breakdown in the zero-row-sum invariant of the softmax Jacobian—a mathematical condition required for accurate backpropagation. The method reconciles E4M3 and E5M2 floating-point format scaling factors across the forward-backward boundary. By adjusting these factors, the process eliminates the stale-delta effect, a systematic gradient bias that compounds errors over training iterations.
Timeline
2024: NAVER Cloud documented irrecoverable FP8 training divergence.
September 29, 2026: The Delta-Matching paper was posted to arXiv.
The Tech Race
The development represents a critical step in overcoming the limitations of the RULER-8K benchmark for 8-bit training. It moves the field beyond the failures documented in 2024, providing a viable path toward utilizing lower-precision formats at scale.
This method aims to lower the compute cost and resource requirements for training future large language models. Developers can expect to see more efficient training workflows once the authors release the code and documented recipes.
The takeaway
Delta-Matching demonstrates that the accuracy gap in 8-bit training is a resolvable mathematical error rather than a hardware limitation. Watch for the forthcoming release of trained checkpoints and data recipes to confirm if these performance gains hold for larger models.
What happens next
The authors plan to release their implementation, trained model checkpoints, and data recipes, which will serve as the next milestone for external validation at scale.
Further reading
For more on the challenges of model training, visit the Artificial Intelligence section.
Source note: This article includes information reported by Tech Times.
Live Poll
Do you believe advancements in AI efficiency will significantly lower the cost of building large models?









