Improving Rare Disease Diagnosis with Selective AI Prediction.

Zhaoyang Jiang, Zhizhong Fu, Yunsoo Kim, Zicheng Li, Xuanqi Peng, Fei Teng, Jiacong Mi, Honghan Wu· August 18, 2026 View original

Key takeaways

  • Standard top-score selective prediction is insufficient for rare disease diagnosis.
  • Small LLMs show very low recall on ultra-rare diseases.
  • The top-two margin can be a better confidence signal for fixed-candidate rankers.
  • Confidence signals must match the specific decision being made.

Who benefits

HealthcarePharmaceuticalsBiotech

Summary

This research explores selective prediction for diagnostic systems, particularly for rare diseases, finding that standard top-score thresholding is insufficient. It highlights that current small LLMs struggle with ultra-rare diseases and that the top-two margin is a better confidence signal for fixed-candidate rankers, though not universally applicable.

Diagnostic AI systems typically rank potential diseases and then decide whether to endorse their top prediction or defer it for human review. This decision is often based on a simple threshold of the highest prediction score. However, this study reveals significant limitations of this approach, especially when dealing with rare diseases. It demonstrates that even small, open-weight Large Language Models (LLMs) achieve very low recall on ultra-rare conditions, indicating that even perfect confidence ranking cannot achieve high selective accuracy in these challenging scenarios. The research emphasizes that the confidence signal used for decision-making must align with the specific decision being made. For systems that provide a fixed list of candidate diseases, the margin between the top two scores proves to be a more effective confidence metric than the top score alone. This margin helps filter out cases where the model is less certain among its top choices. However, the study also proves that relying solely on unlabelled scores cannot definitively determine when switching to the margin will be beneficial, suggesting a nuanced application is required. The findings underscore the need for more sophisticated selective prediction strategies in high-stakes domains like healthcare, particularly for conditions with low prevalence. It suggests that simply improving model accuracy might not be enough; the confidence mechanisms also need to be rethought to ensure reliable and safe AI-assisted diagnosis.

Why it matters

Professionals in healthcare AI development need to understand the limitations of current selective prediction methods, especially for rare conditions, to build more reliable and trustworthy diagnostic tools.

How to implement this in your domain

  1. 1Re-evaluate confidence scoring mechanisms in existing AI diagnostic systems, especially for low-prevalence conditions.
  2. 2Investigate implementing top-two margin-based selective prediction for fixed-candidate ranking systems.
  3. 3Conduct rigorous testing of AI diagnostic tools on diverse datasets, including a significant proportion of rare disease cases.
  4. 4Develop hybrid human-AI workflows where the AI defers uncertain rare disease predictions to human experts.
  5. 5Collaborate with medical professionals to define acceptable recall and precision thresholds for rare disease diagnosis.

Original post by Zhaoyang Jiang, Zhizhong Fu, Yunsoo Kim, Zicheng Li, Xuanqi Peng, Fei Teng, Jiacong Mi, Honghan Wu

"arXiv:2608.14683v1 Announce Type: new Abstract: Given a patient's clinical findings, a diagnostic system ranks possible diseases and must decide when to endorse its first prediction or defer it for review. This decision is usually made by thresholding the top score. Selective pre…"

View on X

Originally posted by Zhaoyang Jiang, Zhizhong Fu, Yunsoo Kim, Zicheng Li, Xuanqi Peng, Fei Teng, Jiacong Mi, Honghan Wu on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses