BACON Improves AI Judge Accuracy with Budgeted Human Calibration.

Lei Shi, Anlan Zhang, Rita Lyu, Zhengmian Hu, Tong Yu, David Arbour, Avi Feller, Saayan Mitra, Ritwik Sinha· July 21, 2026 View original

Summary

This paper introduces BACON, a four-stage pipeline that combines limited human calibration with multiple AI judge outputs to produce more accurate and less biased evaluations. It uses human labels as a calibration anchor, enhancing efficiency and reducing bias in AI-driven model ranking and quality reporting.

Evaluating AI models often relies on AI judges, which are scalable but can introduce biases compared to human preferences. These biases can distort decisions when used for model ranking or quality reporting. The BACON framework addresses this by integrating budgeted human calibration with outputs from multiple AI judges. BACON operates in four stages: it generates comprehensive auxiliary features for each item, including multi-judge scores and contextual embeddings. A small, sampled subset of items receives human labels, which are then used to train a cross-fitted outcome model. This model generates calibrated item-level predictions. These calibrated predictions support two main applications: estimating population-level summary metrics with valid confidence intervals, and providing individual-level surrogate scores for item ranking. By treating AI judges as auxiliary measurements and human labels as the calibration anchor, BACON significantly improves predictive accuracy, ranking consistency, and reduces bias and variance across various tasks and domains, even with limited human annotation budgets.

Why it matters

Professionals relying on AI-driven evaluation systems can achieve more accurate and reliable model assessments, leading to better decision-making in product development and quality assurance. This framework offers a practical way to leverage AI's scalability while mitigating its inherent biases.

How to implement this in your domain

  1. 1Integrate BACON's four-stage pipeline into existing AI model evaluation workflows.
  2. 2Allocate a small budget for human annotators to calibrate AI judge outputs on critical data subsets.
  3. 3Utilize the calibrated item-level predictions for more accurate model ranking and performance reporting.
  4. 4Apply the framework to estimate population-level quality metrics with improved confidence intervals.
  5. 5Train internal teams on the principles of budgeted human calibration for AI evaluation.

Who benefits

TechSoftware DevelopmentAI/ML ConsultingMarket ResearchContent Moderation

Key takeaways

  • AI judges are scalable but prone to biases that distort evaluation outcomes.
  • BACON combines limited human calibration with multiple AI judges for more accurate evaluations.
  • The framework improves predictive accuracy, ranking consistency, and reduces bias and variance.
  • It offers a statistically grounded approach for scalable evaluation with reduced human annotation costs.

Original post by Lei Shi, Anlan Zhang, Rita Lyu, Zhengmian Hu, Tong Yu, David Arbour, Avi Feller, Saayan Mitra, Ritwik Sinha

"arXiv:2607.16239v1 Announce Type: new Abstract: AI judges offer a scalable, low-cost alternative to human evaluation, but their outputs can be biased relative to human preferences and highly item-dependent, varying across judges, tasks, and domains. When uncalibrated AI evaluatio…"

View on X

Originally posted by Lei Shi, Anlan Zhang, Rita Lyu, Zhengmian Hu, Tong Yu, David Arbour, Avi Feller, Saayan Mitra, Ritwik Sinha on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses