MobileForge: New Benchmark for Multi-Screen Mobile App Generation

Fan Wu, Cuiyun Gao, Yiming Huang, Yang Xiao, Yujia Chen, Qing Liao· August 3, 2026 View original

Key takeaways

  • Existing design-to-code benchmarks are insufficient for evaluating multi-screen mobile app generation.
  • MobileForge is the first benchmark for project-level mobile app generation, evaluating build, navigation, visual fidelity, maintainability, and efficiency.
  • Current multimodal LLMs can build and navigate to pages, but interactive navigation and code quality are still lacking.
  • The benchmark introduces improved testing protocols for navigation and visual evaluation.

Who benefits

Software DevelopmentMobile App DevelopmentAI DevelopmentUI/UX Design

Summary

This paper introduces MobileForge, the first benchmark designed to evaluate multimodal LLMs for generating complete, multi-screen mobile applications from visual designs. It addresses limitations of existing benchmarks by focusing on project-level functionality, navigation, and maintainability.

Current benchmarks for design-to-code AI models primarily focus on generating single-page interfaces, which falls short of the complexity required for real-world mobile applications. Real apps demand multiple interconnected screens, shared components, and functional navigation within a buildable codebase. To address this gap, researchers have developed MobileForge, a novel benchmark specifically for project-level multi-screen mobile app generation. MobileForge comprises real mobile apps, human-reviewed screens, structured page-relationship annotations, and navigation test specifications. It enables a comprehensive five-axis evaluation covering build success, navigation, visual fidelity, code maintainability, and efficiency. Initial tests on six leading multimodal LLMs show that while models can generate compilable projects and reach correct pages, interactive navigation remains unreliable, and visual fidelity and code maintainability still require significant improvement. The benchmark also introduces state-isolated navigation testing and an anchor-referenced visual evaluation protocol to enhance reliability.

Why it matters

For developers and product managers in mobile app development, MobileForge provides a crucial tool to accurately assess and drive the progress of AI models capable of generating complex, production-ready mobile applications.

How to implement this in your domain

  1. 1Utilize the MobileForge benchmark to evaluate the performance of current design-to-code LLMs for multi-screen app generation in your organization.
  2. 2Focus AI development efforts on improving interactive navigation, visual fidelity, and code maintainability, identified as weak points by the benchmark.
  3. 3Integrate project-level evaluation metrics, similar to MobileForge's five-axis system, into your internal AI development and testing workflows.
  4. 4Explore the benchmark's methodology for state-isolated navigation testing to enhance the robustness of your app testing strategies.

Original post by Fan Wu, Cuiyun Gao, Yiming Huang, Yang Xiao, Yujia Chen, Qing Liao

"arXiv:2607.28645v1 Announce Type: cross Abstract: Recent multimodal large language models can convert visual designs directly into executable code, but real mobile products require multiple screenshots to become a buildable codebase with shared components and working navigation.…"

View on X

Originally posted by Fan Wu, Cuiyun Gao, Yiming Huang, Yang Xiao, Yujia Chen, Qing Liao on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses