CRATE Improves Mobile Agent Evaluation for Task Completion and Safety

Pengshuai Yang, Zijing Gao, Xue Yu, Benhui Zhuang, Bo Yuan, Junlan Feng· August 24, 2026 View original

Key takeaways

  • CRATE is a VLM-as-judge framework for evaluating mobile agents.
  • It uses step-level reasoning to overcome context overload.
  • CRATE assesses both task completion and operational safety (CRATE-S).
  • The framework significantly improves evaluation accuracy and robustness.

Who benefits

RoboticsAutonomous VehiclesLogisticsSmart ManufacturingAI Development

Summary

Researchers introduce CRATE, a VLM-as-judge framework for automated evaluation of language-guided mobile agents, which performs step-level consequence reasoning and aggregation. CRATE significantly improves task completion assessment and extends to CRATE-S for operational safety, outperforming existing holistic evaluation paradigms.

This research presents CRATE, a novel two-stage VLM-as-judge framework designed for the automated evaluation of language-guided mobile agents. Current evaluation methods often struggle with context overload and primarily focus on task completion, neglecting operational safety. CRATE addresses these limitations by employing a step-level consequence reasoning mechanism. The framework independently extracts task-relevant visual clues and infers action-conditioned state changes at each step. This step-level textual evidence is then aggregated to provide an evidence-grounded evaluation of task completion. Furthermore, CRATE is extended to CRATE-S for assessing operational safety. Experiments demonstrate CRATE's effectiveness and robustness, with CRATE achieving an F1-score of 0.833 on AndroidWorld and CRATE-S reaching 0.697 on MobileRisk, showing strong alignment with ground truths and significant improvement over previous methods.

Why it matters

For professionals developing and deploying mobile AI agents, CRATE offers a more accurate, scalable, and comprehensive evaluation method that considers both task completion and crucial operational safety, leading to more reliable and trustworthy agents.

How to implement this in your domain

  1. 1Adopt CRATE or similar step-level evaluation frameworks for assessing mobile agent performance.
  2. 2Integrate safety assessment (CRATE-S) into the development lifecycle of mobile agents.
  3. 3Utilize VLM-as-judge paradigms to automate and scale agent evaluation processes.
  4. 4Train engineering teams on advanced evaluation techniques for robust agent development.

Original post by Pengshuai Yang, Zijing Gao, Xue Yu, Benhui Zhuang, Bo Yuan, Junlan Feng

"arXiv:2608.20797v1 Announce Type: new Abstract: Evaluating language-guided mobile agents has recently shifted from rule-based to model-based approaches to achieve scalable and automated assessments. However, existing holistic evaluation paradigms process entire trajectories at on…"

View on X

Originally posted by Pengshuai Yang, Zijing Gao, Xue Yu, Benhui Zhuang, Bo Yuan, Junlan Feng on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses

More in AI Engineering & DevTools