New World Models Improve Web Agent Action Selection

Kelvin Li, Dhruv Pendharkar, Anish Pahilajani, Chuyi Shang, Leon Oks, Leonid Karlinsky, Rogerio Feris, Trevor Darrell, Roei Herzig· September 3, 2026 View original

Key takeaways

  • Traditional web agent world models are misaligned with downstream action rankers.
  • Predicted-state matching trains world models to better distinguish between action outcomes.
  • This new training objective improves action ranking and overall task success for web agents.
  • Specialized branching datasets are key to effectively training these discriminative models.

Who benefits

Software DevelopmentE-commerceCustomer ServiceData Science

Summary

This paper introduces "predicted-state matching," a new training objective for world models in web agents that makes predicted states more discriminative. This approach improves action ranking and end-to-end task success for web agents by better distinguishing true outcomes from alternative actions.

Recent advancements in web agents often rely on world models to simulate future web states and select optimal actions. These models typically predict fixed representations like HTML, but this objective doesn't align well with the downstream ranker's need to differentiate between various candidate actions. This research proposes a novel training objective called "predicted-state matching." This new objective ensures that the predicted web state representation can effectively distinguish the actual outcome from states resulting from alternative actions. The models were trained using a specialized branching web-agent dataset, derived from WebArena Go-Browse trajectories, which includes multiple alternative actions and their resulting states for each decision point. Experiments demonstrated that this approach significantly outperforms traditional supervised next-state prediction, leading to improved action ranking and higher end-to-end task success rates on benchmarks like WebPRMBench and WebArena-Lite.

Why it matters

Professionals developing or deploying AI web agents can achieve higher reliability and performance by incorporating these discriminative world models, leading to more effective automation of web-based tasks.

How to implement this in your domain

  1. 1Investigate integrating discriminative world models into existing web automation or agent-based systems.
  2. 2Develop or adapt datasets to include branching trajectories for training world models with predicted-state matching.
  3. 3Benchmark current web agent performance against agents using this new approach for action selection.
  4. 4Collaborate with AI researchers to explore the application of this technique to other interactive AI systems.

Original post by Kelvin Li, Dhruv Pendharkar, Anish Pahilajani, Chuyi Shang, Leon Oks, Leonid Karlinsky, Rogerio Feris, Trevor Darrell, Roei Herzig

"arXiv:2609.02885v1 Announce Type: new Abstract: Recent web agents use world models for test-time action selection by sampling candidate actions, predicting the resulting web states, and ranking them with a ranker model or a Process Reward Model (PRM). These world models are typic…"

View on X

Originally posted by Kelvin Li, Dhruv Pendharkar, Anish Pahilajani, Chuyi Shang, Leon Oks, Leonid Karlinsky, Rogerio Feris, Trevor Darrell, Roei Herzig on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses