New Tool Benchmarks LLM Spreadsheet Creation Abilities

Mansi Uniyal, Agamdeep Singh, Ananya Singha, Priyanshu Gupta, Mukul Singh, Gust Verbruggen, Vu Le, Sumit Gulwani· August 11, 2026 View original

Key takeaways

  • The "workbook time machine" automates the creation of benchmarks for LLM spreadsheet capabilities.
  • `wtmbench` is a new 150-task benchmark for evaluating LLMs on Excel tasks.
  • LLM performance on spreadsheet tasks is heavily influenced by query specificity, agent orchestration, and API choice.
  • This tool helps identify effective AI solutions for automating complex spreadsheet operations.

Who benefits

Financial ServicesBusiness ConsultingData AnalyticsAccountingOffice Automation

Summary

Researchers introduce the "workbook time machine," a pipeline that automatically generates benchmarks for evaluating language models' ability to create derived objects in spreadsheets like formulas, charts, and pivot tables. This tool, applied to public corpora, produced `wtmcorpus` and a curated `wtmbench` for evaluating LLM performance on Excel tasks.

Evaluating the capability of large language models (LLMs) to interact with and manipulate spreadsheets is a complex task. This research presents a novel pipeline called the "workbook time machine," designed to automate the creation of benchmarks for assessing LLMs' proficiency in generating derived spreadsheet objects. These objects include formulas, charts, pivot tables, and conditional formatting, covering a range of complexity. By applying this pipeline to public workbook corpora, the researchers generated `wtmcorpus`, a comprehensive collection of input workbook, output workbook, and query triples. From this larger corpus, they curated `wtmbench`, a focused benchmark comprising 150 tasks with queries at three distinct levels of specificity. Initial evaluations using `wtmbench` on existing spreadsheet manipulation agents and baseline models revealed several critical factors influencing LLM performance. These include the specificity of the query, the orchestration strategy of the agent, and the particular API used to control the spreadsheet. The findings highlight that these elements play a significant role in how effectively LLMs can perform Excel-related tasks.

Why it matters

For professionals looking to automate spreadsheet tasks with LLMs, this benchmark provides a standardized way to evaluate and compare different AI agents, helping to identify the most effective solutions for specific business needs.

How to implement this in your domain

  1. 1Explore `wtmbench` to assess the current capabilities of LLMs or agentic systems for automating spreadsheet tasks relevant to your business.
  2. 2When designing prompts for LLM-driven spreadsheet automation, consider the impact of query specificity on performance.
  3. 3Investigate different agent orchestration strategies and spreadsheet APIs to optimize LLM performance for complex Excel operations.
  4. 4Use the insights from this research to inform the development or selection of AI tools for data analysis and reporting.

Original post by Mansi Uniyal, Agamdeep Singh, Ananya Singha, Priyanshu Gupta, Mukul Singh, Gust Verbruggen, Vu Le, Sumit Gulwani

"arXiv:2608.07873v1 Announce Type: new Abstract: We introduce the workbook time machine, a pipeline that automatically creates benchmarks evaluating the ability of language models to create derived objects in spreadsheets (formulas, charts, pivot tables, and conditional formatting…"

View on X

Originally posted by Mansi Uniyal, Agamdeep Singh, Ananya Singha, Priyanshu Gupta, Mukul Singh, Gust Verbruggen, Vu Le, Sumit Gulwani on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses