xMICD Creates Explainable ICD Code Representations for EHRs

Pat Vatiwutipong, Kumkup Keeratisiwakul, Albert Phuoc Kien Van Truong, Nutcha Yodrabum, Wasin Pansiritanachot, Marvin N. Wright, Thanapon Noraset· August 4, 2026 View original

Key takeaways

  • xMICD offers interpretable, low-dimensional patient representations from ICD codes.
  • It combines clinical groupings with semantic similarity from ICD embeddings.
  • The method achieves high predictive performance comparable to complex embeddings.
  • xMICD enhances the transparency and trustworthiness of clinical machine learning models.

Who benefits

HealthcarePharmaceuticalsHealth InsuranceMedical Research

Summary

xMICD is a new method that generates low-dimensional, interpretable patient representations from ICD codes by combining clinical groupings with embedding similarity. It achieves predictive performance comparable to complex embedding methods while maintaining clinical interpretability for machine learning models.

Effectively representing International Classification of Diseases (ICD) codes from Electronic Health Records (EHRs) is a challenge in clinical risk prediction. Existing methods often force a trade-off: grouping-based representations are interpretable but can lose information, while embedding-based representations offer high predictive power but lack transparency. xMICD (Explainable Representation of Multiple ICD Codes) bridges this gap. It constructs patient representations by integrating clinically meaningful diagnostic groupings with semantic similarity derived from pre-trained ICD embeddings. Instead of simple binary group assignments, xMICD uses similarity-based relative assignments, creating features that show how closely a patient's diagnoses align with various clinical groups. This approach yields predictive performance on par with advanced embedding methods like ICD2Vec, while ensuring that each feature dimension corresponds to an understandable clinical concept, thus enhancing interpretability.

Why it matters

Healthcare professionals and data scientists can build more transparent and trustworthy machine learning models for clinical risk prediction, allowing for better understanding of model decisions and improved patient care.

How to implement this in your domain

  1. 1Adopt xMICD to create interpretable patient representations from ICD codes in your EHR datasets.
  2. 2Integrate clinically meaningful diagnostic groupings with pre-trained ICD embeddings in your feature engineering pipeline.
  3. 3Apply similarity-based relative assignments to generate features that reflect alignment with clinical groups.
  4. 4Benchmark xMICD's predictive performance against existing embedding-based methods on your clinical prediction tasks.
  5. 5Utilize the interpretable features to explain model predictions to clinicians and stakeholders.

Original post by Pat Vatiwutipong, Kumkup Keeratisiwakul, Albert Phuoc Kien Van Truong, Nutcha Yodrabum, Wasin Pansiritanachot, Marvin N. Wright, Thanapon Noraset

"arXiv:2608.00935v1 Announce Type: new Abstract: Electronic Health Records (EHRs) are widely used for clinical risk prediction using machine learning. International Classification of Diseases (ICD) codes provide structured information about patient diagnoses, but representing them…"

View on X

Originally posted by Pat Vatiwutipong, Kumkup Keeratisiwakul, Albert Phuoc Kien Van Truong, Nutcha Yodrabum, Wasin Pansiritanachot, Marvin N. Wright, Thanapon Noraset on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses