New MAGA Framework Improves Cross-Platform GUI Agent Performance

Hang Yan, Zhangxuan GU, Beitong Zhou, Jiaxuan Chen, Runze Li, Yusong Hu, Shuheng Shen, Changhua Meng· August 3, 2026 View original

Key takeaways

  • MAGA improves cross-platform GUI agents by consolidating specialized models.
  • Structured action distillation focuses learning on correcting erroneous actions.
  • A training-only hint mechanism optimizes supervision from expert models.
  • The framework achieves higher success rates compared to existing baselines.

Who benefits

Software DevelopmentUI/UX DesignAutomationCustomer ServiceGaming

Summary

The MAGA framework enhances GUI agents by consolidating specialized models into a single cross-environment policy, using structured action distillation to focus learning on erroneous actions and improve success rates across platforms.

Existing graphical user interface (GUI) agents, often powered by large language models, are typically designed for specific platforms like mobile, web, or desktop. This specialization limits their broader deployment and user experience. The challenge lies in merging these domain-specific experts into a single, versatile agent without corrupting executable actions due to conflicting instructions. Traditional methods like weight merging can lead to inconsistencies, while on-policy distillation often treats all response tokens equally, overlooking the critical role of action tokens in agent-environment interaction. The new MAGA framework addresses this by re-allocating training signals based on the correctness of generated actions. It suppresses invalid distillation signals and prioritizes learning from errors. Additionally, MAGA incorporates a training-only hint mechanism to optimize supervision from specialized teachers without altering the student agent's input. This approach has shown significant improvements, achieving higher mean success rates and outperforming strong baselines across different model scales, demonstrating its effectiveness in creating more robust and adaptable GUI agents.

Why it matters

Professionals developing AI agents for user interfaces can leverage this research to create more versatile and robust agents that operate seamlessly across diverse platforms, reducing development complexity and improving user experience.

How to implement this in your domain

  1. 1Investigate structured action distillation techniques for existing multi-platform agent development.
  2. 2Experiment with re-allocating training signals based on action correctness in your agent training pipelines.
  3. 3Explore incorporating training-only hint mechanisms to refine teacher supervision without altering agent inputs.
  4. 4Evaluate the performance gains of consolidated cross-environment policies compared to domain-specific agents.

Original post by Hang Yan, Zhangxuan GU, Beitong Zhou, Jiaxuan Chen, Runze Li, Yusong Hu, Shuheng Shen, Changhua Meng

"arXiv:2607.29320v1 Announce Type: new Abstract: Graphical user interface (GUI) agents based on large language models are increasingly deployed across mobile, web, and desktop environments. However, existing agents are typically domain-specific, limiting the deployment and user ex…"

View on X

Originally posted by Hang Yan, Zhangxuan GU, Beitong Zhou, Jiaxuan Chen, Runze Li, Yusong Hu, Shuheng Shen, Changhua Meng on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses