DuMateBench Evaluates Autonomous Agents in Real-World Workflows

Zechun Niu, Yukun Zhao, Jiaxin Zhang, Xu Shen, Jinhua Si, Han Tian, Can Xu, Yunfan Song, Jiaxin Mao, Yansong Gao, Yuchen Li, Jianmin Wu, Lingyong Yan, Shuaiqiang Wang, Dawei Yin· August 28, 2026 View original

Key takeaways

  • Existing agent benchmarks don't reflect real-world complexity.
  • DuMateBench uses real user sessions and environmental perturbations.
  • It reveals significant performance gaps in autonomous agents.
  • Robustness depends on both LLM and agent framework capabilities.

Who benefits

Software EngineeringIT OperationsCustomer ServiceRoboticsBusiness Process Automation

Summary

DuMateBench is a new real-session benchmark for evaluating autonomous agents in complex, multi-tool workflows, reconstructed from anonymized user sessions. It exposes agents to real-world environmental complexities like insufficient, unstable, and noisy conditions, revealing substantial performance gaps in strict task completion across various frameworks and LLMs.

Autonomous agents are increasingly being adopted for complex, multi-tool workflows in real-world environments. However, existing benchmarks often separate tasks by application or capability and evaluate agents in cleaner, more stable settings than those encountered in practice. This creates a gap between benchmark performance and real-world utility. To address this, DuMateBench has been introduced as a real-session benchmark. It was reconstructed from anonymized and privacy-screened user sessions from a large-scale production agent platform. Each task in the benchmark retains the relevant pre-solution interaction history, persistent configurations, and workspace state, and is human-verified for accuracy. The benchmark comprises 200 tasks across 8 broad scenarios and 17 fine-grained capability categories, with most tasks requiring coordination of multiple capabilities. These tasks are executed in isolated Docker containers, injected with three types of real-world environmental complexity: Insufficient, Unstable, and Noisy conditions. Performance is assessed using a hybrid deterministic and LLM-as-Judge evaluation protocol. Experiments with five representative autonomous-agent frameworks and four state-of-the-art LLMs revealed significant shortcomings in strict task completion. Further analyses showed that performance under environmental perturbations is influenced by both the LLM's capabilities and the surrounding agent framework. The code and data are publicly available.

Why it matters

This benchmark provides a more realistic and rigorous way to evaluate autonomous agents, helping professionals identify true performance gaps and build more robust, reliable AI systems for real-world deployment.

How to implement this in your domain

  1. 1Utilize DuMateBench to rigorously evaluate the robustness and efficiency of internal autonomous agent deployments.
  2. 2Integrate real-world environmental complexities (insufficient, unstable, noisy conditions) into agent testing protocols.
  3. 3Analyze agent performance under stress to identify specific weaknesses in multi-tool coordination and error handling.
  4. 4Leverage the public dataset and code to contribute to the development of more resilient autonomous agents.

Original post by Zechun Niu, Yukun Zhao, Jiaxin Zhang, Xu Shen, Jinhua Si, Han Tian, Can Xu, Yunfan Song, Jiaxin Mao, Yansong Gao, Yuchen Li, Jianmin Wu, Lingyong Yan, Shuaiqiang Wang, Dawei Yin

"arXiv:2608.26546v1 Announce Type: new Abstract: Autonomous agents are increasingly adopted to complete complex, multi-tool workflows in real-world settings. However, existing benchmarks typically separate tasks by application or capability and evaluate agents in environments that…"

View on X

Primary sources

Originally posted by Zechun Niu, Yukun Zhao, Jiaxin Zhang, Xu Shen, Jinhua Si, Han Tian, Can Xu, Yunfan Song, Jiaxin Mao, Yansong Gao, Yuchen Li, Jianmin Wu, Lingyong Yan, Shuaiqiang Wang, Dawei Yin on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses

More in AI Engineering & DevTools