Review Identifies Gaps in Surgical ML Risk Prediction

Yizhi Dong, Yuhe Ke, Hairil Rizal Abdullah, Yucheng Xing, Kevan Kai Bing Teo, Ling Huang, Mengling Feng· August 3, 2026 View original

Key takeaways

  • Current ML for surgical risk prediction suffers from significant methodological inconsistencies and limitations.
  • Lack of open-access, multi-center datasets hinders reproducibility and generalizability.
  • Incomplete reporting of preprocessing steps and absence of standardized benchmarks are common issues.
  • There is an underutilization of advanced ML techniques like deep learning and multimodal approaches.

Who benefits

HealthcareMedical ResearchHealthTechPharmaceuticals

Summary

A scoping review of 190 studies reveals significant methodological gaps in end-to-end machine learning approaches for surgical risk stratification and outcome prediction using EHR data. Key issues include reliance on single-center datasets, incomplete reporting, lack of standardized benchmarks, and limited use of advanced ML techniques.

Postoperative complications and mortality remain a substantial global health burden, with early identification of high-risk patients being crucial for preventative care. Machine learning (ML) offers a data-driven approach to model complex clinical patterns using electronic health records (EHRs). However, a comprehensive scoping review of 190 studies on ML for surgical risk stratification and outcome prediction has uncovered significant inconsistencies and limitations in current practices.The review highlighted several critical gaps across the ML workflow. Most studies relied on private, single-center datasets, severely limiting reproducibility and generalizability due to a scarcity of open-access surgical datasets. Reporting of essential preprocessing steps, such as handling missing data, feature selection, and class imbalance, was often incomplete.Furthermore, the research found a predominance of conventional ML models and simple neural networks, with deep learning and multimodal approaches being uncommon. The absence of benchmark datasets and standardized evaluation protocols hindered cross-study comparisons. Only about a third of the studies incorporated explainability methods. These findings underscore the need for more rigorous, reproducible, and clinically meaningful ML development in perioperative care.

Why it matters

Healthcare professionals and AI developers can use this review to understand the current state and critical shortcomings of ML in surgical risk prediction, guiding future research, development, and implementation towards more robust and clinically useful tools.

How to implement this in your domain

  1. 1Advocate for and contribute to the creation of open-access, multi-center surgical datasets to improve ML model generalizability.
  2. 2Implement standardized reporting guidelines for ML methodology in healthcare, covering data preprocessing, model selection, and evaluation.
  3. 3Prioritize the development and adoption of benchmark datasets and evaluation protocols for surgical risk prediction.
  4. 4Integrate explainability methods into ML models for clinical applications to foster trust and facilitate clinical interpretation.
  5. 5Explore deep learning and multimodal approaches for surgical risk stratification, moving beyond conventional ML models.

Original post by Yizhi Dong, Yuhe Ke, Hairil Rizal Abdullah, Yucheng Xing, Kevan Kai Bing Teo, Ling Huang, Mengling Feng

"arXiv:2607.29090v1 Announce Type: new Abstract: Postoperative adverse events, including mortality and morbidity, remain a major global burden, many of which are preventable through early identification of high-risk patients and targeted perioperative care. Accurate risk stratific…"

View on X

Originally posted by Yizhi Dong, Yuhe Ke, Hairil Rizal Abdullah, Yucheng Xing, Kevan Kai Bing Teo, Ling Huang, Mengling Feng on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses