Comparing Feature Selection for Opioid Use Disorder Prediction

Zihan Ding, Yinan Liu, Tengfei Ma, Rachel Wong, George Leibowitz, Benjamin Littenberg, Xia Zheng, Richard N. Rosenthal, Fusheng Wang· August 6, 2026 View original

Key takeaways

  • Effective feature selection is crucial for accurate and interpretable EHR-based OUD prediction.
  • NTK sensitivity offers the best balance of accuracy and stability among tested methods.
  • LLM-guided selection provides complementary, clinically meaningful insights.
  • Performance gains diminish beyond a moderate number of selected features.

Who benefits

HealthcarePharmaceuticalsPublic HealthHealthTechInsurance

Summary

This study compares five feature selection methods for predicting opioid use disorder (OUD) using EHR diagnosis codes, evaluating their impact on predictive performance, stability, and representation of infrequent codes. NTK sensitivity emerged as the best overall method, balancing accuracy and stability.

Predicting Opioid Use Disorder (OUD) from Electronic Health Records (EHR) is challenging due to the high dimensionality, sparsity, noise, and redundancy of diagnosis-related features. This research conducted a comparative analysis of five distinct feature selection methods: recurrence enrichment, NTK-motivated early gradient sensitivity, LightGBM-SHAP, Elastic Net, and large language model (LLM)-guided semantic selection. The goal was to assess how each method impacts predictive performance, resampling stability, and the ability to represent less common diagnosis codes. The study found that increasing the number of features generally improved performance, though with diminishing returns beyond a certain point. Among the methods tested, NTK sensitivity demonstrated the most favorable balance between predictive accuracy and stability. Interestingly, while LLM-guided semantic selection did not achieve the highest standalone performance, it provided valuable, clinically meaningful signals that complemented the other methods, highlighting the potential for hybrid approaches in medical predictive modeling.

Why it matters

Improving feature selection for EHR data can lead to more accurate and interpretable predictive models for conditions like OUD, aiding early intervention and personalized patient care while reducing computational burden.

How to implement this in your domain

  1. 1Apply NTK-motivated early gradient sensitivity for feature selection in new EHR-based predictive models.
  2. 2Explore combining LLM-guided semantic selection with other methods to enhance clinical interpretability.
  3. 3Benchmark different feature selection techniques on specific healthcare prediction tasks within your organization.
  4. 4Develop a standardized preprocessing and evaluation framework for comparing feature selection methods.

Original post by Zihan Ding, Yinan Liu, Tengfei Ma, Rachel Wong, George Leibowitz, Benjamin Littenberg, Xia Zheng, Richard N. Rosenthal, Fusheng Wang

"arXiv:2608.04180v1 Announce Type: new Abstract: Feature selection is a critical step in electronic health record (EHR)-based predictive modeling, where input variables are often high-dimensional, sparse, noisy, and redundant. Large feature sets not only increase computational bur…"

View on X

Originally posted by Zihan Ding, Yinan Liu, Tengfei Ma, Rachel Wong, George Leibowitz, Benjamin Littenberg, Xia Zheng, Richard N. Rosenthal, Fusheng Wang on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses