EarlyDx Benchmark Evaluates LLM Diagnostic Accuracy at Hospital Admission

Jiahui Li, Ruili Fang, Zishuai Liu, Yutong Guo, Nan Yang, Wenzhan Song, Jin Lu, Fei Dou· August 3, 2026 View original

Key takeaways

  • EarlyDx is a new benchmark for evaluating LLM diagnostic accuracy at hospital admission.
  • Current LLMs struggle to synthesize admission-time evidence reliably, especially for inferred diagnoses.
  • A significant gap exists between LLM and clinician performance in early diagnosis.
  • The benchmark emphasizes open-ended, evidence-supported diagnosis from limited data.

Who benefits

HealthcareAI DevelopmentMedical TechnologyEmergency Services

Summary

Researchers introduce EarlyDx, a new benchmark for open-ended, evidence-supported early diagnosis at hospital admission, built from 154,834 emergency department encounters. It reveals that current LLMs struggle to reliably synthesize admission-time evidence, particularly for inferred diagnoses, highlighting a significant gap compared to clinician performance.

A new research paper introduces EarlyDx, a large-scale benchmark designed to evaluate the ability of large language models (LLMs) to generate open-ended, evidence-supported diagnoses at the critical point of hospital admission. Unlike previous benchmarks that rely on closed code sets or discharge diagnoses, EarlyDx uses 154,834 emergency department encounters from MIMIC-IV, restricting evidence to what's available at admission and supervising with diagnoses recorded during the ED encounter. A key feature of EarlyDx is its LLM auditor, which verifies every free-text label for support from the admission-time evidence, scoring only fully supported diagnoses. Using a semantic LLM-as-judge protocol, the evaluation revealed that no tested system—whether general, medical-specialized, or in-domain post-trained—could reliably synthesize admission-time evidence. Zero-shot models primarily extracted information, recovering only a small fraction of diagnoses requiring inference. While post-training improved inference-dependent recall, a substantial gap remains, and no system achieved a clinician's balance of sensitivity and precision for time-critical conditions. The full construction and evaluation pipeline are publicly released.

Why it matters

For healthcare AI developers, EarlyDx provides a crucial, realistic benchmark to assess and improve LLM performance in early, time-sensitive clinical diagnosis, which is vital for patient outcomes and operational efficiency.

How to implement this in your domain

  1. 1Utilize the EarlyDx benchmark to rigorously evaluate and improve LLM performance for early diagnostic support in healthcare.
  2. 2Focus LLM training on enhancing inference capabilities from incomplete clinical histories, not just extraction.
  3. 3Develop AI systems that can balance diagnostic sensitivity and precision for time-critical medical conditions.
  4. 4Collaborate with clinicians to understand the nuances of admission-time diagnostic reasoning for AI model refinement.

Original post by Jiahui Li, Ruili Fang, Zishuai Liu, Yutong Guo, Nan Yang, Wenzhan Song, Jin Lu, Fei Dou

"arXiv:2607.28788v1 Announce Type: new Abstract: Clinical diagnosis at hospital admission must be made rapidly from limited, incomplete evidence. Existing diagnosis-prediction benchmarks are poorly suited to this setting: they restrict prediction to closed code sets, exclude free-…"

View on X

Originally posted by Jiahui Li, Ruili Fang, Zishuai Liu, Yutong Guo, Nan Yang, Wenzhan Song, Jin Lu, Fei Dou on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses