Advanced Optimizers Boost Large LLM Pretraining Efficiency

Mikail Khona, Aditya Vavre, Boxiang Wang, Deyu Fu, Hao Wu, Mike Chrzanowski, Bryan Catanzaro, Dheevatsa Mudigere, Jeff Pool, Michael Lightstone, Mohammad Shoeybi, Mostofa Patwary, Nima Tajbakhsh, Tijmen Blankevoort· July 24, 2026 View original

Summary

This work adapts and enhances higher-order optimizers like SOAP and Muon to overcome stability and computational challenges in large-scale LLM pretraining, demonstrating their superior performance over AdamW at multi-billion-parameter scales and large batch sizes. The researchers also introduce a layer-wise distributed optimizer and system-level improvements, releasing a codebase for emerging optimization algorithms.

Researchers have made significant strides in adapting and enhancing higher-order optimizers, such as SOAP and Muon, to improve the efficiency and stability of large-scale Large Language Model (LLM) pretraining. While these optimizers offer faster convergence than traditional methods like AdamW, their adoption has been limited by computational costs and numerical stability issues at scale. The team identified and addressed instabilities in SOAP at large batch sizes by proposing algorithmic modifications, including per-step QR orthogonalization and improved preconditioning, which successfully eliminated loss spikes and enabled stable training. A comprehensive empirical study comparing SOAP, Muon, and AdamW on multi-billion-parameter models trained on trillions of tokens revealed that SOAP and Muon consistently outperformed AdamW. Notably, these advanced optimizers maintained training stability and quality even at batch sizes up to 100 million tokens for next-token prediction, where AdamW's performance degraded. To facilitate efficient training at such scales, the researchers developed a layer-wise distributed optimizer compatible with Megatron-LM, balancing memory and communication while preserving the optimizers' convergence benefits. They also implemented system-level improvements and released their codebase to support the research community.

Why it matters

For organizations developing and training large language models, these advancements offer a path to significantly reduce training time and computational resources, enabling the creation of more powerful and cost-effective AI systems.

How to implement this in your domain

  1. 1Evaluate integrating SOAP or Muon optimizers into your LLM pretraining workflows.
  2. 2Experiment with the proposed algorithmic modifications for improved stability at large batch sizes.
  3. 3Explore the layer-wise distributed optimizer for efficient scaling of LLM training.
  4. 4Leverage the released codebase to accelerate research and development in advanced optimization.

Who benefits

AI/ML DevelopmentCloud ComputingResearch & DevelopmentSoftware Development

Key takeaways

  • Higher-order optimizers like SOAP and Muon outperform AdamW for large LLM pretraining.
  • Algorithmic modifications improve stability at large batch sizes, preventing loss spikes.
  • A new layer-wise distributed optimizer enables efficient scaling of these methods.
  • These advancements can significantly reduce LLM training time and cost.

Original post by Mikail Khona, Aditya Vavre, Boxiang Wang, Deyu Fu, Hao Wu, Mike Chrzanowski, Bryan Catanzaro, Dheevatsa Mudigere, Jeff Pool, Michael Lightstone, Mohammad Shoeybi, Mostofa Patwary, Nima Tajbakhsh, Tijmen Blankevoort

"arXiv:2607.20548v1 Announce Type: new Abstract: Higher-order optimizers such as Muon and SOAP offer faster convergence than AdamW, but their computational cost and numerical stability challenges have limited adoption at scale. In this work, we adapt and enhance preconditioned gra…"

View on X

Originally posted by Mikail Khona, Aditya Vavre, Boxiang Wang, Deyu Fu, Hao Wu, Mike Chrzanowski, Bryan Catanzaro, Dheevatsa Mudigere, Jeff Pool, Michael Lightstone, Mohammad Shoeybi, Mostofa Patwary, Nima Tajbakhsh, Tijmen Blankevoort on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses