Safe-Psych Benchmark Reveals LLMs Struggle with Diagnostic Uncertainty

Oriana Presacan, Andreea Grama, Larisa Irimin\u{a}, Alireza Nik, Jaya Ojha, Vajira Thambawita, Ciprian I. B\u{a}cil\u{a}, Bogdan Ionescu, Michael A. Riegler· July 16, 2026 View original

Key takeaways

  • LLMs struggle to recognize and appropriately handle incomplete clinical information, often diagnosing prematurely.
  • Current safety prompting may shift errors from premature diagnosis to excessive abstention, not true uncertainty handling.
  • Models rarely proactively seek clarification unless explicitly prompted, a critical flaw in clinical decision support.
  • New benchmarks like Safe-Psych are vital for evaluating LLM safety and calibration in dynamic healthcare settings.

Who benefits

HealthcareAI DevelopmentMedical TechnologyPharmaceuticalsClinical Research

Summary

Safe-Psych is a new sequential evaluation benchmark for LLMs in psychiatry, designed to test how models handle evolving diagnostic uncertainty. It reveals that even strong LLMs frequently diagnose prematurely or abstain excessively when information is incomplete, rarely seeking clarification unless explicitly prompted.

This paper introduces Safe-Psych, a novel sequential evaluation benchmark specifically designed to assess how large language models (LLMs) perform under conditions of evolving diagnostic uncertainty in clinical psychiatry. While LLMs are increasingly used for healthcare decision support, existing benchmarks often assume complete information is available upfront, failing to test a model's ability to request clarification or abstain when data is insufficient. Safe-Psych comprises over 1,000 real-world psychiatric clinical notes, segmented to simulate the incremental disclosure of evidence. At each stage, psychiatrist-derived action labels (DIAGNOSE, CLARIFY, or ABSTAIN) indicate the appropriate response. The evaluation of state-of-the-art LLMs using this benchmark revealed a significant limitation: even highly capable models struggle with incomplete clinical information. They often exhibit "under-abstention," prematurely diagnosing in over 60% of cases, or shift to "excessive abstention" with safety-aware prompting. The sequential evaluation highlighted that models frequently diagnose before sufficient evidence is gathered and rarely proactively seek clarification. These premature diagnoses were also less accurate. The core issue identified across evaluated models is their difficulty in recognizing when clinical evidence is incomplete and more information is needed. Safe-Psych is released to foster research into improving LLM safety in healthcare.

Why it matters

For professionals in healthcare AI, this research underscores a critical safety gap in current LLMs: their inability to recognize and appropriately handle diagnostic uncertainty. This has profound implications for deploying AI in clinical settings, emphasizing the need for models that can "know what they don't know" and ask for more information.

How to implement this in your domain

  1. 1Integrate uncertainty quantification and clarification-seeking mechanisms into LLM-based diagnostic tools.
  2. 2Develop training methodologies that explicitly teach LLMs to identify and respond to incomplete information in clinical contexts.
  3. 3Utilize benchmarks like Safe-Psych to rigorously evaluate the safety and reliability of healthcare AI systems under evolving information conditions.
  4. 4Design user interfaces for clinical AI that prompt for additional information when the model indicates uncertainty, rather than presenting a definitive answer.

Original post by Oriana Presacan, Andreea Grama, Larisa Irimin\u{a}, Alireza Nik, Jaya Ojha, Vajira Thambawita, Ciprian I. B\u{a}cil\u{a}, Bogdan Ionescu, Michael A. Riegler

"arXiv:2607.13036v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for decision support in healthcare, but clinical evidence is often incomplete or evolving. When the available information is insufficient to support a reliable answer, models shou…"

View on X

Originally posted by Oriana Presacan, Andreea Grama, Larisa Irimin\u{a}, Alireza Nik, Jaya Ojha, Vajira Thambawita, Ciprian I. B\u{a}cil\u{a}, Bogdan Ionescu, Michael A. Riegler on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses