New Benchmark Reveals AI Agents Struggle with Data Science Workflows

Mizanur Rahman, Mohammed Saidul Islam, Ridwan Mahbub, Md Tahmid Rahman Laskar, Shafiq Joty, Enamul Hoque Prince· August 12, 2026 View original

Key takeaways

  • DSAgentBench evaluates AI agents on end-to-end data science workflows.
  • Current agents struggle significantly with real-world data science tasks.
  • Top proprietary agents achieve moderate success, open-source agents perform poorly.
  • Key failure points include tool orchestration and multi-step reasoning.

Who benefits

TechConsultingFinanceResearchData Analytics

Summary

DSAgentBench, a new benchmark, evaluates AI agents' ability to automate end-to-end data science workflows in real computer environments. Results show a significant gap, with even top agents achieving only 56.70% success and open-source models below 1%, highlighting failures in tool orchestration and multi-step reasoning.

A new benchmark called DSAgentBench has been introduced to rigorously test the capabilities of AI agents in automating complete data science workflows within realistic computing environments. Unlike previous benchmarks, DSAgentBench requires agents to interact with various tools like notebooks, IDEs, terminals, and databases, mirroring the complexity of real-world data science practice. It features 275 diverse tasks covering the entire data science lifecycle, from wrangling to visualization and validation. The evaluation criteria are comprehensive, verifying analytical correctness, visual outputs, and model performance, not just code execution. Extensive testing with 15 different AI models, both proprietary and open-source, revealed a substantial performance gap. The strongest agent, Claude-4.6-Sonnet, achieved only 56.70% task success, while open-source agents performed below 1%. Common failure points included tool orchestration, operating system grounding, and multi-step reasoning, indicating that current agentic systems are far from fully automating complex data science tasks.

Why it matters

Professionals relying on or developing AI agents for data science tasks need to understand their current limitations. This benchmark provides a realistic assessment, guiding expectations and future development efforts towards truly autonomous data science.

How to implement this in your domain

  1. 1Assess current AI agent capabilities against real-world data science workflow requirements.
  2. 2Prioritize agent development efforts on improving tool orchestration and multi-step reasoning.
  3. 3Integrate robust error handling and debugging mechanisms into agentic systems.
  4. 4Invest in training agents on diverse, complex, and multi-tool interaction scenarios.

Original post by Mizanur Rahman, Mohammed Saidul Islam, Ridwan Mahbub, Md Tahmid Rahman Laskar, Shafiq Joty, Enamul Hoque Prince

"arXiv:2608.10366v1 Announce Type: new Abstract: Real-world data science involves long-horizon workflows that span data wrangling, exploration, modeling, visualization, and validation, and require coordinated use of tools such as notebooks, IDEs, terminals, browsers, and databases…"

View on X

Originally posted by Mizanur Rahman, Mohammed Saidul Islam, Ridwan Mahbub, Md Tahmid Rahman Laskar, Shafiq Joty, Enamul Hoque Prince on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses