Who&When Pro Benchmarks LLM Failure Attribution in AI Agents.
Key takeaways
- Who&When Pro is a new, large-scale benchmark for automated failure attribution in AI agents.
- It provides 12,326 labeled failed trajectories across diverse scenarios.
- The research reveals systematic patterns in LLM failure attribution.
- This benchmark offers empirical guidance for developing better attribution systems.
Who benefits
Summary
This research introduces Who&When Pro, a large-scale benchmark for automated failure attribution in agentic AI systems, featuring 12,326 failed trajectories with golden labels across diverse modalities and scenarios. It reveals systematic patterns in how LLMs attribute failures, offering empirical guidance for future attribution systems.
Why it matters
For professionals developing, deploying, or managing AI agents, accurate failure attribution is critical for debugging, improving system reliability, and accelerating development cycles. This benchmark provides tools and insights to achieve that.
How to implement this in your domain
- 1Utilize the Who&When Pro benchmark to evaluate the failure attribution capabilities of your current LLM-based agent systems.
- 2Analyze the systematic patterns identified in the research to inform the design of your own attribution mechanisms.
- 3Integrate automated failure attribution tools into your AI agent development and monitoring pipelines.
- 4Experiment with different LLM families and protocols for attribution to find the most effective approach for your specific agents.
- 5Contribute to the development of more robust failure attribution systems by leveraging insights from this benchmark.
Original post by Jiale Liu, Huajun Xi, Shaokun Zhang, Yifan Zeng, Tianwei Yue, Chi Wang, Jian Kang, Qingyun Wu, Huazheng Wang
"arXiv:2607.09996v1 Announce Type: new Abstract: Automated failure attribution uses LLMs to identify where and why agentic systems fail. As agents become more capable, their failures become subtler, making automated attribution increasingly important. We introduce Who&When Pro, a…"
View on XOriginally posted by Jiale Liu, Huajun Xi, Shaokun Zhang, Yifan Zeng, Tianwei Yue, Chi Wang, Jian Kang, Qingyun Wu, Huazheng Wang on X · view source
Want to go deeper?
Turn these trends into skills with Learnijoy's hands-on AI & tech courses.
Explore coursesMore in AI Engineering & DevTools
Zapier vs. Tray: Enterprise Automation Platform Comparison for 2026
This post compares Zapier and Tray.io, evaluating which platform is better suited for enterprise automation needs by balancing power and ease of use. It argues that the best tools scale for complex requirements while remaining intuitive for all users.
Cross-Regime Bayesian Optimization Boosts Algorithmic Trading Signals
This paper introduces a cross-regime Bayesian optimization approach for hyperparameter selection in algorithmic trading, targeting robustness across different market regimes. It finds that a hybrid ensemble of XGBoost and TabNet achieves an annualized return of 51.26% and a Sharpe ratio of 2.44, outperforming individual models and demonstrating significant out-of-sample generalization.
Emotional Preferences Regulate Goal Priorities in Reinforcement Learning Agents
This paper proposes a computational framework where higher-level goals autonomously generate state-dependent emotional preferences to regulate the priorities of competing lower-level objectives in reinforcement learning agents. It demonstrates how this emergent preference function exhibits contextual priority switching and improves performance over fixed-preference strategies in multi-objective exploration environments.