AI M&M Framework Proposed for Clinical AI Failure Review.

Paulius Mui, Dean F. Sittig, Steve Labkoff, Sanjay Basu· September 2, 2026 View original

Key takeaways

  • Existing safety mechanisms are insufficient for learning from individual clinical AI failures.
  • AI M&M is a structured, blameless framework for reviewing AI-related errors and near-misses.
  • It classifies events by Trigger, Mechanism, Clinical Pathway, and Corrective Action.
  • The framework aims to convert individual failures into actionable institutional learning.

Who benefits

HealthcareMedical DevicesAI/ML DevelopmentRegulatory AffairsInsurance

Summary

A new framework, AI Morbidity and Mortality (AI M&M), is proposed for structured, blameless review of clinical AI failures and near-misses. It aims to explain how risks emerge from AI system interactions with clinicians and workflows, providing actionable institutional learning beyond aggregate monitoring or traditional safety reports.

As clinical artificial intelligence becomes more integrated into healthcare, existing safety mechanisms are proving inadequate for understanding and learning from individual AI-related errors. Traditional aggregate model monitoring identifies performance shifts, and patient safety reporting captures adverse events, but neither explains the complex interplay between AI systems, clinicians, workflows, and institutional controls that leads to risk. To address this gap, researchers propose the AI Morbidity and Mortality (AI M&M) framework. This structured, blameless approach facilitates case-based review of clinical AI failures. It includes standardized case intake, evidence preservation, investigator-level reconstruction, tool-in-loop attribution, and corrective-action tracking. Events are classified across four dimensions: Trigger, Mechanism, Clinical Pathway, and Corrective Action, enabling a clear understanding of how vulnerabilities are exposed, risks produced, care impacted, and remediation assigned. The framework was demonstrated with five illustrative cases, showing high agreement among reviewers.

Why it matters

This framework provides a critical tool for healthcare organizations to systematically learn from AI failures, enhancing patient safety and improving the responsible deployment of AI in clinical settings.

How to implement this in your domain

  1. 1Evaluate existing incident reporting systems for their suitability in capturing AI-specific failure modes.
  2. 2Pilot the AI M&M framework in a specific clinical department using AI-powered tools.
  3. 3Establish a multidisciplinary team to conduct blameless reviews of AI-related incidents.
  4. 4Develop clear protocols for evidence preservation and reconstruction for AI system failures.
  5. 5Integrate lessons learned from AI M&M reviews into AI system design, deployment, and clinician training.

Original post by Paulius Mui, Dean F. Sittig, Steve Labkoff, Sanjay Basu

"arXiv:2609.00076v1 Announce Type: new Abstract: Clinical artificial intelligence is increasingly embedded in real-world care, yet existing safety mechanisms are poorly suited to reconstructing and learning from individual AI-related errors and near-misses. Aggregate model monitor…"

View on X

Originally posted by Paulius Mui, Dean F. Sittig, Steve Labkoff, Sanjay Basu on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses

More in AI News & Tools

AI Engineering & DevToolsAI News & Tools

CRAFT Enables Explainable AI for 6G RAN Networks

CRAFT is a data-centric method that generates verified (input, trace, label) datasets to fine-tune Small Language Models (SLMs) for pre-hoc explainability in AI-native 6G RAN. It overcomes the cold-start barrier of RL methods, achieving high accuracy and F1 scores with significantly less energy consumption.

Pranshav Gajjar, Vijay K ShahSep 2, 2026
AI Engineering & DevToolsAI ResearchAI News & Tools

Safin-1 Enhances AI Safety via Memory-Native State Evolution.

Safin-1 is a new family of foundation models that achieves "Safety from Within" by integrating safety-relevant capabilities through memory routing and state evolution, allowing for test-time adaptation of persistent capability states without modifying the backbone. This reframes memory as an active substrate for evolving model behavior.

Ming Zhang, Kaisen Yang, Shu Yu, Ermo Hua, Zhekai Chen, Cheng Jin, Jingnan Zheng, Yi Zhang, Zhongtian Ma, Jiawei Zhou, Sirui Chen, Qiaosheng Zhang, Xiang Wang, Ning Ding, Xia Hu, Bowen Zhou, Youbang Sun, Chaochao LuSep 2, 2026
AI InvestingAI Engineering & DevToolsAI News & Tools

Foundation Models' Role in Electricity Price Forecasting Examined.

This study compares nine foundation models against market-specific benchmarks for electricity price forecasting and battery arbitrage, finding that while TabPFN models statistically outperform, their economic dominance is not universal and depends on risk tolerance and bidding strategies. Foundation models cannot fully replace specialized models.

Arkadiusz Lipiecki, Rafa{\l} WeronSep 2, 2026