LLMs Struggle with Statistical Problem Formulation, New Benchmark Reveals

Chen Wang, Junzhe Zhao, Xin Cong, Wanlu Deng, Ke Deng· September 3, 2026 View original

Key takeaways

  • LLMs currently struggle with the crucial task of statistical problem formulation.
  • A new benchmark, StatFormBench, evaluates LLMs on problem classification and variable identification.
  • Even top models show moderate accuracy, indicating a need for significant improvement.
  • Human oversight remains essential for accurate statistical problem setup when using LLMs.

Who benefits

Data ScienceAI/ML EngineeringConsultingResearch & DevelopmentEducation

Summary

This research introduces StatFormBench, a new benchmark to evaluate LLMs' ability to formulate statistical problems from informal goals and diverse data. It finds that even the best models achieve only moderate accuracy in classifying problem types and identifying relevant variables, highlighting a significant gap in their statistical reasoning capabilities.

Large Language Models (LLMs) are increasingly used as assistants in data science, but their ability to handle the crucial initial step of statistical problem formulation has been largely unexamined. This upstream task involves interpreting informal user goals and heterogeneous data to identify the appropriate statistical task and relevant variables. To address this, researchers formalized statistical problem formulation into two subtasks: classification of the statistical problem and identification/role assignment of variables. They then developed StatFormBench, a comprehensive benchmark comprising 1,013 samples from statistics textbooks and data science cases, covering 20 coarse-grained and 85 fine-grained problem categories. Testing 14 LLMs (both open and closed-source) on StatFormBench revealed significant limitations. The best zero-shot models achieved only 72.0% fine-grained classification accuracy and 63.2% variable set overlap. No single model consistently excelled across both subtasks, and advanced prompting strategies offered only marginal improvements. This indicates a substantial gap in LLMs' ability to translate real-world scenarios into structured statistical problems.

Why it matters

For data scientists and AI professionals, this research highlights a critical limitation of current LLMs: their struggle with the nuanced, upstream task of problem formulation. Relying solely on LLMs for this step could lead to incorrect analyses or missed insights, necessitating human oversight and improved model capabilities.

How to implement this in your domain

  1. 1Integrate human oversight and validation steps when using LLMs for initial statistical problem formulation in data science workflows.
  2. 2Develop fine-tuning datasets specifically tailored to improve LLMs' understanding of statistical problem types and variable identification.
  3. 3Design interactive tools that allow data scientists to easily correct or refine LLM-generated problem formulations and variable assignments.
  4. 4Explore multi-modal approaches that combine LLMs with structured knowledge bases or expert systems for more robust statistical problem formulation.

Original post by Chen Wang, Junzhe Zhao, Xin Cong, Wanlu Deng, Ke Deng

"arXiv:2609.01982v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as assistants for statistical and data science work, yet existing evaluations largely assume the analysis target is already specified. In practice, users arrive with informal goals…"

View on X

Originally posted by Chen Wang, Junzhe Zhao, Xin Cong, Wanlu Deng, Ke Deng on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses