Wuying-Browser-Agent Sets New Standard for Long-Horizon Browser Automation

AIMAE Team, Tianxiang Chen, Yan Cheng, Zhangye Han, Xiaowei Li, Chang Liu, Cheng Liu, Zhongqiang Ma, Long Peng, Xiaobing Tu, Yinggui Wang, Hongliang Wei, Chen Wu, Daiping Xin, Kunyu Zhou, Pengyang Zhou, Peiyuan Chen, Ziyuan Chen, Yutao Deng, Chunyu Dong, Xiangyu Fu, Yicheng Feng, Ruian He, Haochen Li, Miancan Liu, Zhengqin Liu, Wei Peng, Jinkui Ren, Haoyu Tan, Dong Xiao, Rongkun Xue, Shujian Yang, Xianhang Ye, Ziqi Yuan, Ziyang Yu, Linghan Zhang, Xiantao Zhang, Xuanpu Zhao, Yinan Zhao, Zhenghui Zhao, Bin Zhu, Likai Zou· August 19, 2026 View original

Key takeaways

  • Wuying-Browser-Agent sets a new state of the art for long-horizon, real-world browser automation.
  • It addresses challenges like error recovery and complex UI navigation through a unified framework.
  • The framework includes structured execution, specialized SFT, and divergence-aware online RL.
  • BrowserBench is a new benchmark for evaluating agents on complex, multi-step web tasks.

Who benefits

Business Process AutomationSoftware TestingE-commerceData ScrapingCustomer Service

Summary

Wuying-Browser-Agent is a new unified framework that establishes a new open-source state of the art for long-horizon browser agents, achieving high performance on complex real-world web tasks. It addresses challenges like sustained decision-making, error recovery, and complex UI navigation through structured execution, specialized SFT, and divergence-aware online reinforcement learning.

While browser agents have shown promise on simple, short demonstrations, their performance in real-world deployments often falls short due to the need for sustained decision-making, error recovery, and navigation of complex user interfaces over long horizons. Researchers argue that addressing this gap requires alignment across the entire AI pipeline, not just scaling up models. They introduce Wuying-Browser-Agent, a comprehensive framework designed to tackle these real-world challenges. The framework incorporates a structured browser harness for stable execution and context management, along with Reflection and UI-specialized Curriculum SFT (RUIC-SFT) to explicitly train agents on recovery trajectories and complex UI interactions. Furthermore, Divergence-Aware Online GRPO (DAO-GRPO) is used to improve long-horizon credit assignment through potential-based reward shaping and divergence-aware step weighting. To provide a more realistic evaluation, the team also developed BrowserBench, a bilingual real-web benchmark comprising 350 tasks averaging 37.9 steps, specifically designed to expose long-horizon failure modes. Wuying-Browser-Agent-27B achieved impressive results, reaching 80.6% on WebVoyager, 66.7% on Online-Mind2Web, and 65.1% on BrowserBench, setting a new open-source state of the art. The pipeline's general agentic ability was also demonstrated across other benchmarks.

Why it matters

For professionals seeking to automate complex web-based workflows, Wuying-Browser-Agent offers a significant leap forward in reliability and capability, enabling AI agents to handle real-world browser tasks that require sustained interaction, error recovery, and adaptability.

How to implement this in your domain

  1. 1Explore integrating Wuying-Browser-Agent or its underlying techniques into your organization's web automation projects.
  2. 2Develop internal benchmarks similar to BrowserBench to rigorously test the long-horizon capabilities of your browser agents.
  3. 3Implement structured browser harnesses and context management for more stable agent execution.
  4. 4Investigate specialized fine-tuning methods (like RUIC-SFT) to improve agent recovery from errors and navigation of complex UIs.
  5. 5Apply advanced reinforcement learning techniques (like DAO-GRPO) for better credit assignment in long-duration tasks.

Original post by AIMAE Team, Tianxiang Chen, Yan Cheng, Zhangye Han, Xiaowei Li, Chang Liu, Cheng Liu, Zhongqiang Ma, Long Peng, Xiaobing Tu, Yinggui Wang, Hongliang Wei, Chen Wu, Daiping Xin, Kunyu Zhou, Pengyang Zhou, Peiyuan Chen, Ziyuan Chen, Yutao Deng, Chunyu Dong, Xiangyu Fu, Yicheng Feng, Ruian He, Haochen Li, Miancan Liu, Zhengqin Liu, Wei Peng, Jinkui Ren, Haoyu Tan, Dong Xiao, Rongkun Xue, Shujian Yang, Xianhang Ye, Ziqi Yuan, Ziyang Yu, Linghan Zhang, Xiantao Zhang, Xuanpu Zhao, Yinan Zhao, Zhenghui Zhao, Bin Zhu, Likai Zou

"arXiv:2608.17319v1 Announce Type: new Abstract: Browser agents perform well on short, clean demonstrations, but real deployment is fundamentally different: agents must sustain dozens of decisions on live websites while recovering from mistakes and navigating complex UIs. We argue…"

View on X

Originally posted by AIMAE Team, Tianxiang Chen, Yan Cheng, Zhangye Han, Xiaowei Li, Chang Liu, Cheng Liu, Zhongqiang Ma, Long Peng, Xiaobing Tu, Yinggui Wang, Hongliang Wei, Chen Wu, Daiping Xin, Kunyu Zhou, Pengyang Zhou, Peiyuan Chen, Ziyuan Chen, Yutao Deng, Chunyu Dong, Xiangyu Fu, Yicheng Feng, Ruian He, Haochen Li, Miancan Liu, Zhengqin Liu, Wei Peng, Jinkui Ren, Haoyu Tan, Dong Xiao, Rongkun Xue, Shujian Yang, Xianhang Ye, Ziqi Yuan, Ziyang Yu, Linghan Zhang, Xiantao Zhang, Xuanpu Zhao, Yinan Zhao, Zhenghui Zhao, Bin Zhu, Likai Zou on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses

More in AI Engineering & DevTools