Dion3 Optimizer Accelerates Orthogonal Updates for AI Models

Noah Amsel, Jack Zhang, Kwangjun Ahn, Ali Naeimi, Austin Feng, Berlin Chen, Tri Dao, John Langford· August 13, 2026 View original

Key takeaways

  • Dion3 significantly reduces the overhead of orthogonal updates in AI optimizers.
  • It achieves up to 6x faster step times compared to Muon while maintaining performance.
  • Innovations include a new Gram Newton-Schulz algorithm and a fractional orthogonalization rule.
  • Dion3 is available as a drop-in replacement, enhancing efficiency for large-scale model training.

Who benefits

AI/ML ResearchCloud ComputingHigh-Performance ComputingData CentersAutonomous Systems

Summary

Dion3 is a revised optimizer that significantly reduces the computational and communication overhead of orthogonal updates, a key component of the Muon optimizer. It achieves up to 6x faster step times while matching or improving model loss, making large-scale AI training more efficient.

The Muon optimizer, known for its orthogonal updates, often incurs substantial overhead due to its cubic-time Newton-Schulz orthogonalization step. This computational burden is further compounded by communication overhead when model weights are sharded across multiple devices, limiting Muon's practical benefits. Dion3 is introduced as a comprehensive revision designed to tackle this overhead at every level of the software stack. It features a new Gram Newton-Schulz algorithm that reduces FLOP costs, specialized CuteDSL kernels that exploit symmetry for acceleration, and a megabatching strategy to minimize communication overhead. A crucial innovation in Dion3 is a simplified update rule that orthogonalizes only a fraction of the momentum matrix's rows at each step, further cutting costs. This approach not only improves upon previous "compressed" versions like Dion in both speed and performance but also achieves up to a 6x reduction in optimizer step time while maintaining or even improving the loss achieved by the original Muon. Dion3 is available as a drop-in replacement via the dion package.

Why it matters

AI engineers and researchers can leverage Dion3 to significantly accelerate the training of large-scale models that benefit from orthogonal updates, leading to faster experimentation, reduced infrastructure costs, and more efficient development cycles.

How to implement this in your domain

  1. 1Integrate the Dion3 optimizer into your deep learning training pipelines as a drop-in replacement for Muon.
  2. 2Benchmark Dion3's performance against existing optimizers on your specific large-scale models.
  3. 3Experiment with the fraction of momentum matrix rows orthogonalized to find the optimal balance between speed and performance.
  4. 4Utilize Dion3's multi-GPU capabilities to optimize distributed training of large AI models.

Original post by Noah Amsel, Jack Zhang, Kwangjun Ahn, Ali Naeimi, Austin Feng, Berlin Chen, Tri Dao, John Langford

"arXiv:2608.11612v1 Announce Type: new Abstract: The Muon optimizer incurs a significant overhead cost due to its cubic-time Newton-Schulz orthogonalization step. When weights are sharded, communication overhead compounds this computational cost, eroding the benefits of Muon in ma…"

View on X

Originally posted by Noah Amsel, Jack Zhang, Kwangjun Ahn, Ali Naeimi, Austin Feng, Berlin Chen, Tri Dao, John Langford on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses