Muon Optimizer Scales Efficiently for Large Diffusion Transformers.

Chenghao Li, Xiao Han, Xinxin Huang, Wei Liu, Boyang Li, Bing Xiao, Heran Zhang, Juanma Perez Rua, Ke Xu, Kangning Liu, Linjun Kuang, Na Li, Tan Wang, Tian Xie, Wei Peng, Yang Pei, Yifan Xu, Yuanhao Zhai, Yuwei Lin, Zhe Wang, Zihao He, Daniel Li, Junbiao Tang, Ziyang Jiang, Dake Chen· August 24, 2026 View original

Key takeaways

  • Periodic Row-wise Muon significantly improves the training efficiency of large Diffusion Transformers.
  • It maintains superior generative quality compared to AdamW while drastically reducing training time and communication.
  • The method addresses the computational overheads of the original Muon optimizer at scale.
  • Optimized distributed implementation is key to realizing its full efficiency benefits.

Who benefits

AI DevelopmentCloud ComputingMedia & EntertainmentResearch & AcademiaAutomotive

Summary

Researchers introduce Periodic Row-wise Muon, an optimized version of the Muon optimizer, which significantly improves training efficiency and generative quality for large Diffusion Transformers (DiTs) by reducing computational and communication overhead. This new approach maintains Muon's benefits while drastically cutting training time and communication volume.

The Muon optimizer has shown promise in improving the training of large models by balancing updates across singular directions, leading to better generative quality. However, its original implementation faced significant computational and communication overheads, particularly when scaled to very large Diffusion Transformers (DiTs) ranging from 1.3B to 15B parameters. This overhead stemmed from its 5-step Newton-Schulz iteration and full-momentum materialization at every optimization step. To address these challenges, a new method called Periodic Row-wise Muon has been developed. This approach performs a full spectral update less frequently, specifically once every K steps, and uses a more efficient, low-cost row-wise constrained update in between. The design also includes a co-designed distributed implementation that works directly with sharded momentum and accelerates spectral refreshes through optimized communication. Experimental results demonstrate that Periodic Row-wise Muon preserves the original Muon's superior generative quality, achieving 12.9-19.1% better quality than AdamW. Crucially, it reduces optimizer time by 46.9-54.3%, end-to-end step time by 15.7-24.3%, and logical communication volume by 66.7%, leading to 33.7-64.8% less active training time to reach optimal generative quality.

Why it matters

This research offers a significant advancement in optimizing the training of large AI models, enabling faster development cycles and more efficient resource utilization for high-quality generative AI applications. Professionals can achieve better model performance with reduced computational costs.

How to implement this in your domain

  1. 1Evaluate current Diffusion Transformer training pipelines for optimizer bottlenecks.
  2. 2Investigate integrating Periodic Row-wise Muon into existing large-scale generative model training frameworks.
  3. 3Benchmark the performance and efficiency gains against current AdamW or vanilla Muon implementations.
  4. 4Optimize distributed training setups to leverage the co-designed communication-computation overlap features.
  5. 5Consider contributing to or adopting open-source implementations of this optimized optimizer.

Original post by Chenghao Li, Xiao Han, Xinxin Huang, Wei Liu, Boyang Li, Bing Xiao, Heran Zhang, Juanma Perez Rua, Ke Xu, Kangning Liu, Linjun Kuang, Na Li, Tan Wang, Tian Xie, Wei Peng, Yang Pei, Yifan Xu, Yuanhao Zhai, Yuwei Lin, Zhe Wang, Zihao He, Daniel Li, Junbiao Tang, Ziyang Jiang, Dake Chen

"arXiv:2608.20818v1 Announce Type: new Abstract: The matrix-aware optimizer Muon improves large model training by balancing updates across singular directions, yet its scaling behavior and end-to-end efficiency on large Diffusion Transformers (DiTs) remain unclear. We first establ…"

View on X

Originally posted by Chenghao Li, Xiao Han, Xinxin Huang, Wei Liu, Boyang Li, Bing Xiao, Heran Zhang, Juanma Perez Rua, Ke Xu, Kangning Liu, Linjun Kuang, Na Li, Tan Wang, Tian Xie, Wei Peng, Yang Pei, Yifan Xu, Yuanhao Zhai, Yuwei Lin, Zhe Wang, Zihao He, Daniel Li, Junbiao Tang, Ziyang Jiang, Dake Chen on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses

More in AI Engineering & DevTools