DataSpace Benchmarks Data Agents for Heterogeneous Analytics

Boyan Li, Zhuowen Liang, Yupeng Xie, Xiaotian Lin, Tianqi Luo, Xinyu Liu, Yizhang Zhu, Zhangyang Peng, Yuan Li, Zhengxuan Zhang, Jiayi Zhang, Nan Tang, Guoliang Li, Yuyu Luo· August 5, 2026 View original

Key takeaways

  • DataSpace is a new benchmark for evaluating data agents on heterogeneous data analytics tasks.
  • Current frontier multimodal models struggle with integrating diverse data sources and performing joins.
  • The benchmark highlights the need for improved reliability and performance in data agents.
  • Harness choice significantly impacts agent accuracy, indicating architectural importance.

Who benefits

Data ScienceBusiness IntelligenceFinancial ServicesHealthcareConsulting

Summary

DataSpace is a new benchmark designed to evaluate data agents' ability to perform verifiable analytics across diverse data sources like databases, files, and multimedia. It features 410 cross-language tasks and 7,439 artifacts, revealing that current frontier models struggle with multimodal evidence integration and joins, with the best accuracy reaching only 66.34%.

Data agents are emerging as crucial tools for natural-language analytics across complex organizational workspaces, where data is often scattered across various formats such as databases, structured files, documents, and multimedia. Existing benchmarks typically focus on isolated aspects like structured querying or retrieval, failing to address the challenges of heterogeneous evidence discovery and verifiable tabular output. To bridge this gap, researchers introduced DataSpace, a comprehensive benchmark for data agents. DataSpace comprises 410 cross-language tasks and over 7,400 artifacts, totaling 15.01 GB, spanning CSV, JSON, SQLite, Markdown, PDF, and video formats. Agents are tasked with producing complete, verifiable tabular results from a given question and workspace. The benchmark, which also served as the KDD Cup 2026 evaluation, revealed significant challenges for current AI models. Even the best of six frontier multimodal models achieved only 66.34% accuracy, with multimodal evidence integration and joins consistently reducing performance. The choice of agent harness also created a substantial 15.36-point accuracy spread, indicating that DataSpace remains largely unsaturated and highlights critical areas for improving data agent reliability.

Why it matters

DataSpace provides a much-needed, rigorous benchmark for evaluating and advancing data agents, which are essential for automating complex data analysis and decision-making in real-world enterprise environments.

How to implement this in your domain

  1. 1Review the DataSpace benchmark to understand the current limitations of data agents in heterogeneous environments.
  2. 2Evaluate existing internal data analysis workflows to identify areas where data agents could provide value but currently struggle.
  3. 3Prioritize research and development efforts on improving multimodal evidence integration and complex data joins for agentic systems.
  4. 4Experiment with different agent harnesses and backbone models to optimize performance on diverse data tasks.
  5. 5Contribute to the DataSpace benchmark or create similar internal benchmarks to drive agent development tailored to specific organizational needs.

Original post by Boyan Li, Zhuowen Liang, Yupeng Xie, Xiaotian Lin, Tianqi Luo, Xinyu Liu, Yizhang Zhu, Zhangyang Peng, Yuan Li, Zhengxuan Zhang, Jiayi Zhang, Nan Tang, Guoliang Li, Yuyu Luo

"arXiv:2608.03451v1 Announce Type: new Abstract: Data agents enable natural-language analytics over organizational workspaces, where relevant evidence may be scattered across databases, structured files, long documents, and multimedia. Existing benchmarks largely isolate structure…"

View on X

Originally posted by Boyan Li, Zhuowen Liang, Yupeng Xie, Xiaotian Lin, Tianqi Luo, Xinyu Liu, Yizhang Zhu, Zhangyang Peng, Yuan Li, Zhengxuan Zhang, Jiayi Zhang, Nan Tang, Guoliang Li, Yuyu Luo on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses