LEMUR 2 Benchmark Unlocks Neural Network Diversity for AI Design

Tolgay Atinc Uzun, Waleed Khalid, Saif U Din, Sai Revanth Mulukuledu, Akashdeep Singh, Chandini Vysyaraju, Raghuvir Duvvuri, Avi Goyal, Yashkumar Rajeshbhai Lukhi, Muhammad A. Hussain, Krunal Jesani, Usha Shrestha, Yash Mittal, Roman Kochnev, Pritam Kadam, Mohsin Ikram, Harsh R. Moradiya, Alice Arslanian, Dmitry Ignatov, Radu Timofte· July 9, 2026 View original

Key takeaways

  • LEMUR 2 is a comprehensive benchmark for neural network diversity, covering generation, evaluation, and deployment.
  • It includes over 14,000 architectures and 750,000 training records from various generation methods, including LLM-guided synthesis.
  • The framework provides real-device performance data for mobile and VR platforms, crucial for practical deployment.
  • LEMUR 2 supports cross-domain analysis and advances LLM-driven AutoML and architectural generalization.

Who benefits

AI DevelopmentSoftware EngineeringMobile GamingVirtual RealityAutomotive

Summary

LEMUR 2 is a new large-scale, extensible framework that unifies generative, evaluative, and deployment pipelines for neural networks, featuring over 14,000 distinct architectures and 750,000 training records. It supports LLM-driven AutoML and cross-platform evaluation, providing a rich dataset for advancing AI design and architectural generalization.

Existing Neural Architecture Search (NAS) benchmarks often provide a limited view of the architectural design space, focusing on narrow, task-specific applications and lacking comprehensive cross-domain or deployment-aware evaluations. Addressing these limitations, LEMUR 2 introduces a groundbreaking, large-scale framework designed to foster neural network diversity. This framework integrates generative, evaluative, and deployment pipelines, offering a holistic approach to AI design. LEMUR 2 encompasses an extensive collection of over 14,000 unique neural network architectures and more than 750,000 structured training records. These records detail model performance, hyperparameters, and task outcomes. The diverse architectures were generated using various advanced methods, including AST-based code mutation, genetic and reinforcement learning evolution, fractal architecture generation, and synthesis guided by Large Language Models (LLMs). Beyond generation and training, LEMUR 2 incorporates NN-VR and NN-Lite pipelines for automated deployment and latency benchmarking across heterogeneous mobile and Unity-based VR platforms. This provides crucial real-device performance metadata. The benchmark covers multimodal tasks, image captioning, text-to-image synthesis, and language modeling, enabling in-depth analysis of architectural transferability across different domains. By linking diverse architectures, tasks, and deployment data, LEMUR 2 establishes a new foundation for reproducible, data-driven AI design, significantly advancing the emerging paradigm of LLM-driven AutoML and architectural generalization across modalities and hardware.

Why it matters

For AI researchers and engineers, LEMUR 2 provides an unprecedented resource for exploring, evaluating, and deploying diverse neural network architectures. It accelerates the development of more efficient, robust, and generalizable AI models, especially in the context of LLM-driven AutoML and cross-platform deployment.

How to implement this in your domain

  1. 1Utilize the LEMUR 2 dataset to benchmark novel neural network architectures against a wide range of existing designs and tasks.
  2. 2Integrate LEMUR 2's deployment pipelines (NN-VR, NN-Lite) to evaluate model performance and latency on real-world mobile and VR platforms.
  3. 3Leverage the diverse architectural generation methods within LEMUR 2 to inspire or guide your own Neural Architecture Search efforts.
  4. 4Explore the dataset for fine-tuning LLMs to generate or optimize neural network architectures, advancing LLM-driven AutoML.
  5. 5Conduct cross-domain analysis using LEMUR 2 to understand architectural transferability and generalization across different AI tasks.

Original post by Tolgay Atinc Uzun, Waleed Khalid, Saif U Din, Sai Revanth Mulukuledu, Akashdeep Singh, Chandini Vysyaraju, Raghuvir Duvvuri, Avi Goyal, Yashkumar Rajeshbhai Lukhi, Muhammad A. Hussain, Krunal Jesani, Usha Shrestha, Yash Mittal, Roman Kochnev, Pritam Kadam, Mohsin Ikram, Harsh R. Moradiya, Alice Arslanian, Dmitry Ignatov, Radu Timofte

"arXiv:2607.06839v1 Announce Type: new Abstract: Existing NAS benchmarks (e.g., NAS-Bench, NATS-Bench) cover only narrow, task-specific regions of the architectural design space and lack cross-domain or deployment-aware evaluation. LEMUR 2 introduces a large-scale, extensible fram…"

View on X

Originally posted by Tolgay Atinc Uzun, Waleed Khalid, Saif U Din, Sai Revanth Mulukuledu, Akashdeep Singh, Chandini Vysyaraju, Raghuvir Duvvuri, Avi Goyal, Yashkumar Rajeshbhai Lukhi, Muhammad A. Hussain, Krunal Jesani, Usha Shrestha, Yash Mittal, Roman Kochnev, Pritam Kadam, Mohsin Ikram, Harsh R. Moradiya, Alice Arslanian, Dmitry Ignatov, Radu Timofte on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses