Auto-Fill Uses Specialist LLMs for Accurate Missing Data Prediction

Yurong Liu, Yeye He, Haoyu Dong, Junjie Xing, Shi Han, Dongmei Zhang, Surajit Chaudhuri· July 23, 2026 View original

Summary

Auto-Fill is a new approach that uses an ensemble of three specialist small language models (SLMs) to accurately predict missing values in tabular data. It combines world knowledge, text-based, and code-based reasoning with a calibrated abstention mechanism, achieving superior accuracy and cost efficiency compared to state-of-the-art reasoning models.

This paper introduces Auto-Fill, a novel approach for accurately predicting missing cell values in tabular data, a critical task in data cleaning. While existing state-of-the-art reasoning models show promise, they are often costly to deploy at scale and prone to overconfidence, leading to hallucinations or false positives. Auto-Fill addresses these limitations by recognizing that high-precision missing-value prediction requires a blend of world knowledge, text-based reasoning, and code-based reasoning. The proposed system post-trains three specialist small language models (SLMs), each optimized for one of these distinct capabilities. A calibrated ensemble mechanism dynamically selects the most confident specialist or abstains from prediction, ensuring high accuracy and reliability. Extensive experiments across 11 benchmarks and 2200 real tables demonstrate that Auto-Fill achieves superior accuracy compared to frontier models like o3-pro, Gemini 3 Pro, and DeepSeek R1, all while operating at less than 1% of their cost. This highlights the effectiveness of specialization and calibrated abstention in tabular data imputation.

Why it matters

Data professionals, analysts, and engineers can significantly improve data quality and reduce manual effort in data cleaning by using Auto-Fill for highly accurate and cost-effective missing value prediction.

How to implement this in your domain

  1. 1Download and integrate the Auto-Fill framework into your data preprocessing pipelines for tabular data.
  2. 2Evaluate Auto-Fill's performance on your specific datasets with missing values, comparing it to existing imputation methods.
  3. 3Leverage the specialist SLMs for tasks requiring world knowledge, text-based, or code-based reasoning for data completion.
  4. 4Utilize the calibrated ensemble mechanism to ensure high-precision predictions and manage abstention for uncertain cases.

Who benefits

Data AnalyticsBusiness IntelligenceHealthcareFinanceE-commerce

Key takeaways

  • Auto-Fill uses specialist SLMs for highly accurate missing value prediction in tabular data.
  • It combines world knowledge, text, and code-based reasoning.
  • A calibrated ensemble mechanism ensures high precision and allows for abstention.
  • Auto-Fill outperforms state-of-the-art models at a fraction of the cost.

Original post by Yurong Liu, Yeye He, Haoyu Dong, Junjie Xing, Shi Han, Dongmei Zhang, Surajit Chaudhuri

"arXiv:2607.19847v1 Announce Type: new Abstract: Predicting missing cell values in tabular data is a fundamental problem in data cleaning. While state-of-the-art reasoning models show great promise in predicting missing values in tables, by reasoning holistically across rows and c…"

View on X

Originally posted by Yurong Liu, Yeye He, Haoyu Dong, Junjie Xing, Shi Han, Dongmei Zhang, Surajit Chaudhuri on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses