SearchAuditor Diagnoses Failures in Long-Horizon Search Agents

Zhixiang Liang, Yifei Liu, Yidan Huang, Haozhe Zhao, Beichen Huang, Jiaqi Wang, Nan Duan, Qiong Cao· August 7, 2026 View original

Key takeaways

  • Long-horizon search agents are prone to complex, propagating reasoning errors.
  • SearchAuditor is a framework for localizing, attributing, and repairing these failures.
  • It uses evidence-grounded adjudication and outperforms frontier LLM baselines.
  • The framework significantly improves agent recovery from errors, enhancing reliability.

Who benefits

AI DevelopmentSoftware EngineeringCybersecurityCustomer ServiceResearch & Development

Summary

Researchers introduce SearchAuditor, a multi-perspective auditing framework designed to localize, attribute, and repair failures in long-horizon web search agents. It uses evidence-grounded adjudication to outperform even frontier models in diagnosing complex reasoning errors in search trajectories.

Deep search agents, which interact with the web over long horizons to answer complex questions, are prone to subtle reasoning errors that can propagate and lead to incorrect answers. Diagnosing these failures is challenging due to the length and complexity of execution traces, often exceeding human capacity for manual inspection. To address this, a new benchmark, SearchAuditBench, and a framework, SearchAuditor, have been developed. SearchAuditBench comprises 1,243 failed trajectories from various open-weight models across five deep-search benchmarks, each expertly annotated with error steps, root causes, and repair suggestions. Building on this, SearchAuditor is a multi-perspective auditing framework that effectively localizes, attributes, and repairs these search-agent failures through evidence-grounded adjudication. Experimental results show that even advanced baseline models, when powered by frontier LLMs like GPT-5.5, achieve only a 26.6% end-to-end pass rate. In contrast, SearchAuditor consistently outperforms all baselines, achieving a 32.3% end-to-end pass rate. Furthermore, its repairs enable agents to better recover from errors, highlighting its significant potential for improving the reliability of long-horizon search agents.

Why it matters

As AI agents become more autonomous and interact with complex environments like the web, robust debugging and failure attribution tools are essential for their reliability and trustworthiness. SearchAuditor provides a critical capability for professionals developing and deploying such agents, enabling them to diagnose and fix errors more efficiently.

How to implement this in your domain

  1. 1Integrate SearchAuditor into your development pipeline for long-horizon AI agents to automatically detect and diagnose failures.
  2. 2Utilize the framework's localization and attribution capabilities to pinpoint root causes of agent errors.
  3. 3Implement the suggested repair mechanisms from SearchAuditor to improve agent robustness and recovery from failures.
  4. 4Contribute to or leverage benchmarks like SearchAuditBench to rigorously test and validate your agent's performance.
  5. 5Establish a feedback loop where human experts review SearchAuditor's diagnoses and repairs to continuously improve the system.

Original post by Zhixiang Liang, Yifei Liu, Yidan Huang, Haozhe Zhao, Beichen Huang, Jiaqi Wang, Nan Duan, Qiong Cao

"arXiv:2608.05212v1 Announce Type: new Abstract: Deep search agents tackle challenging questions through long-horizon web interactions, a process that is both complex and fragile: small reasoning errors may propagate through long, noisy trajectories into fluent but incorrect answe…"

View on X

Originally posted by Zhixiang Liang, Yifei Liu, Yidan Huang, Haozhe Zhao, Beichen Huang, Jiaqi Wang, Nan Duan, Qiong Cao on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses