Agentic AI Framework Boosts Glaucoma Detection Accuracy

Jalil Jalili, Hossein Taghizad, Anuwat Jiravarnsirikul, Christopher Bowd, Akram Belghith, Raheleh Kafieh, Christopher A. Girkin, Sally L. Baxter, Robert N. Weinreb, Linda M. Zangwill, Mark Christopher· August 11, 2026 View original

Key takeaways

  • LLMs alone have limitations in medical image interpretation, including hallucination and inconsistency.
  • An agentic AI framework combines LLMs with specialized deep learning tools.
  • This framework significantly improves glaucoma detection accuracy and consistency.
  • It suggests a future for orchestrated multi-agent systems in medical AI.

Who benefits

HealthcareMedical DevicesHealthTechPharmaceuticals

Summary

This research introduces an agentic AI framework that integrates large language models (LLMs) with specialized deep learning tools to overcome LLM limitations like hallucination, low accuracy, and inconsistency in glaucoma detection from fundus photography. The framework significantly improved classification accuracy, reduced cup-to-disc ratio error, and enhanced run-to-run consistency, matching specialist accuracy in some configurations.

Large language models (LLMs) show potential in medical image interpretation but are often hampered by issues such as generating incorrect information (hallucination), limited accuracy, and inconsistent results across different runs. These limitations are particularly critical in high-stakes medical applications like disease detection. Researchers have developed and validated an agentic AI framework specifically for glaucoma detection from fundus photography, designed to mitigate these LLM shortcomings. The framework operates in three steps: an initial LLM assessment, followed by function calls to specialized deep learning tools for image quality assessment, glaucoma classification, and optic disc/cup segmentation, and finally, an LLM reflection phase that integrates the initial impression with the tool outputs. This agentic workflow dramatically improved classification accuracy by 16 to 47 percentage points, bringing it within 6 points of a fellowship-trained glaucoma specialist, and even matching specialist accuracy on one dataset. It corrected LLM-alone issues like positive bias and stochastic variability, significantly reducing cup-to-disc ratio error and boosting run-to-run consistency. This suggests a shift towards orchestrated multi-agent systems in medical AI, combining LLM reasoning with specialized, reliable tools.

Why it matters

For healthcare professionals, medical AI developers, and health tech companies, this framework offers a robust solution to deploy AI for diagnostic tasks, overcoming critical LLM limitations to achieve higher accuracy and reliability in patient care.

How to implement this in your domain

  1. 1Identify medical image interpretation tasks where LLMs show promise but suffer from accuracy or consistency issues.
  2. 2Design a multi-step agentic workflow that combines LLM reasoning with specialized, validated deep learning tools.
  3. 3Integrate function calling capabilities within your LLM pipeline to invoke specific tools for tasks like image quality assessment or segmentation.
  4. 4Implement an LLM reflection stage to synthesize initial assessments with outputs from specialized tools for a final decision.
  5. 5Validate the agentic framework against expert human performance and existing LLM-alone approaches using relevant medical datasets.

Original post by Jalil Jalili, Hossein Taghizad, Anuwat Jiravarnsirikul, Christopher Bowd, Akram Belghith, Raheleh Kafieh, Christopher A. Girkin, Sally L. Baxter, Robert N. Weinreb, Linda M. Zangwill, Mark Christopher

"arXiv:2608.07651v1 Announce Type: new Abstract: Large language models (LLMs) show promise in medical image interpretation but suffer from hallucination, limited accuracy, and run-to-run inconsistency. We developed and validated an agentic AI framework integrating LLMs with specia…"

View on X

Originally posted by Jalil Jalili, Hossein Taghizad, Anuwat Jiravarnsirikul, Christopher Bowd, Akram Belghith, Raheleh Kafieh, Christopher A. Girkin, Sally L. Baxter, Robert N. Weinreb, Linda M. Zangwill, Mark Christopher on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses