WANDR Benchmark Evaluates AI Research Agents on Complex Tasks

Vitaliy Polshkov, Marcin Pitera, Jeremy Yang, Kirill Priemko, Maksim Gaiduk, Aleksandr Nikolenko, Denis Bykov, Clare Southern, Denis Yarats, Jerry Ma· August 18, 2026 View original

Key takeaways

  • WANDR is a new benchmark for evaluating AI research agents on complex data-collection tasks.
  • It features 500 realistic tasks requiring both broad discovery and deep investigation.
  • The benchmark uses dynamic judges to verify records against current facts.
  • Current AI systems perform poorly on WANDR, indicating significant room for improvement.

Who benefits

Market ResearchConsultingFinancial ServicesLegalData Analytics

Summary

WANDR (Wide ANd Deep Research) is a new benchmark featuring 500 realistic, challenging data-collection tasks designed to evaluate AI research agents on their ability to discover entities, investigate them through multiple web searches, and return verifiable records with sources. The benchmark uses task-specific judges for evaluation, revealing that current production systems are far from saturating its challenges, particularly with increasing target volume and hierarchy depth.

A new benchmark called WANDR (Wide ANd Deep Research) has been introduced to rigorously assess the capabilities of AI research agents. This benchmark comprises 500 complex data-collection tasks that mirror real-world scenarios, demanding agents to both broadly discover entities and deeply investigate each one using coordinated web searches. The ultimate goal is to produce independently verifiable records, complete with supporting sources and excerpts. WANDR's tasks are structured as hierarchical qualification keys, specifying the number of entities, relationships, and evidence required at each level. This design supports diverse professional workflows such as market mapping, due diligence, and talent sourcing, with target record counts ranging from dozens to thousands. Unlike benchmarks with static answer sets, WANDR employs dynamic, task-specific judges that refetch cited pages to verify records against current facts. Initial evaluations of six production research systems on WANDR indicate that the benchmark is highly challenging. The strongest system achieved only 0.363 soft F1 and 0.133 hard F1, with performance significantly degrading as the volume and depth of the hierarchy increased. Key bottlenecks identified include incomplete discovery, missing enrichment, and inadequate evidence construction, suggesting ample room for improvement in current AI research agents.

Why it matters

Professionals developing or utilizing AI agents for information gathering, market research, or due diligence can use WANDR to rigorously evaluate and improve their systems' ability to perform complex, verifiable research tasks.

How to implement this in your domain

  1. 1Access the WANDR benchmark and evaluation harness from the provided GitHub repository.
  2. 2Integrate WANDR tasks into your AI agent development and testing pipeline.
  3. 3Benchmark your existing AI research systems against WANDR to identify performance gaps.
  4. 4Analyze the specific failure modes (e.g., incomplete discovery, missing evidence) to guide agent improvements.
  5. 5Contribute to the community-governed benchmark by submitting new tasks or improvements.

Original post by Vitaliy Polshkov, Marcin Pitera, Jeremy Yang, Kirill Priemko, Maksim Gaiduk, Aleksandr Nikolenko, Denis Bykov, Clare Southern, Denis Yarats, Jerry Ma

"arXiv:2608.14747v1 Announce Type: new Abstract: WANDR (Wide ANd Deep Research) is a benchmark of 500 realistic, challenging data-collection tasks for research agents. Each task requires a system to discover a large set of entities that satisfy specified criteria (breadth), invest…"

View on X

Originally posted by Vitaliy Polshkov, Marcin Pitera, Jeremy Yang, Kirill Priemko, Maksim Gaiduk, Aleksandr Nikolenko, Denis Bykov, Clare Southern, Denis Yarats, Jerry Ma on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses