Multi-LLM Systems Improve Reliability with Uncertainty-Aware Trust Estimation.

Jiawei Zheng, Jiazhen Zhang· July 24, 2026 View original

Summary

This research introduces a method for aggregating predictions from multiple Large Language Models (LLMs) by estimating their trustworthiness based on the quality of their probabilistic predictions, especially under varying reliability. It adapts structured expert judgment to penalize overconfident incorrect predictions and favor well-calibrated LLMs.

Current methods for combining predictions from multiple Large Language Models often assume all models are equally reliable, which can be problematic when dealing with diverse or unreliable LLMs. This new approach addresses this by formulating multi-LLM aggregation as an uncertainty-aware trust estimation problem. It borrows from decision theory's structured expert judgment, using context-aware calibration questions to assess an expert LLM's reliability based on its probabilistic prediction quality. The method employs Cooke-style log weighting, which specifically penalizes overconfident incorrect predictions and rewards experts that are well-calibrated. Evaluations across various LLM panels, including homogeneous, heterogeneous, and contaminated setups, demonstrate that this Cooke weighting significantly improves accuracy and reliability, especially when dealing with diverse or unreliable LLMs. This suggests that effective multi-LLM aggregation requires not just combining outputs, but carefully calibrating trust based on each model's uncertainty.

Why it matters

Professionals building or deploying LLM-powered applications can achieve more robust and reliable systems by intelligently combining outputs from multiple models, especially when dealing with diverse or potentially unreliable LLMs.

How to implement this in your domain

  1. 1Evaluate current multi-LLM aggregation strategies for potential vulnerabilities to unreliable expert models.
  2. 2Investigate integrating uncertainty-aware trust estimation techniques, such as Cooke-style log weighting, into existing LLM ensemble architectures.
  3. 3Develop or adapt calibration questions to assess the probabilistic prediction quality of individual LLMs within an ensemble.
  4. 4Monitor the performance of aggregated LLM systems, paying close attention to accuracy-reliability balance, particularly in heterogeneous or adversarial scenarios.

Who benefits

Software DevelopmentAI/ML ConsultingFinancial ServicesHealthcareCybersecurity

Key takeaways

  • Naive aggregation of LLM predictions can be vulnerable to unreliable or adversarial models.
  • Uncertainty-aware trust estimation improves multi-LLM system reliability and accuracy.
  • Cooke-style log weighting effectively penalizes overconfident incorrect predictions.
  • Calibrating trust is crucial for robust LLM ensembles, especially with heterogeneous models.

Original post by Jiawei Zheng, Jiazhen Zhang

"arXiv:2607.20529v1 Announce Type: new Abstract: Large Language Model (LLM) ensembles are increasingly used to improve reliability by combining predictions from multiple LLMs. However, existing aggregation methods typically assume that all models are equally trustworthy, overlooki…"

View on X

Originally posted by Jiawei Zheng, Jiazhen Zhang on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses