Multi-Omics ML Models Predict Breast Cancer ER Status

Priyanka Paudel, Madan Baduwal· July 21, 2026 View original

Summary

A systematic benchmarking study evaluated classical machine learning models for Estrogen Receptor (ER) status prediction in breast cancer using multi-omics data (transcriptomic, genomic, proteomic). Random Forest achieved the best performance, demonstrating that RNA expression is the strongest predictor and multi-omic integration offers modest but consistent improvements.

Predicting Estrogen Receptor (ER) status is crucial for breast cancer diagnosis, prognosis, and treatment. This study conducted a rigorous benchmarking analysis of various classical machine learning models to predict ER status using multi-omics data from the TCGA-BRCA cohort. The datasets included transcriptomic (RNA expression), genomic (copy number variation), and proteomic (RPPA) information. The research employed a robust experimental framework, incorporating stratified train-test splitting, cross-validation, class imbalance handling, and fold-specific feature selection to ensure reliable evaluation and prevent data leakage. Models such as Random Forest, XGBoost, LightGBM, SVM, and Logistic Regression were tested in both single-omic and multi-omic configurations. Findings indicated that RNA expression data provided the most potent predictive signal. While multi-omic integration yielded only modest improvements, these gains were consistent across models. Random Forest emerged as the top performer in the integrated multi-omic setting, achieving a balanced accuracy of 90.3% and an ROC-AUC of 97.1%. The recurrent selection of known biologically relevant genes further validated the models' findings.

Why it matters

This research provides a validated framework and identifies effective machine learning approaches for a critical breast cancer biomarker prediction, potentially aiding in more precise diagnosis and personalized treatment strategies.

How to implement this in your domain

  1. 1Explore multi-omics data integration strategies for predictive modeling in other disease areas.
  2. 2Adopt rigorous validation frameworks, including stratified splitting and class imbalance handling, for medical ML projects.
  3. 3Consider Random Forest as a strong baseline model for high-dimensional biological datasets.
  4. 4Collaborate with medical professionals to translate predictive model insights into clinical decision support tools.

Who benefits

HealthcarePharmaceuticalsBiotechnologyMedical DiagnosticsResearch & Development

Key takeaways

  • Multi-omics data can effectively predict breast cancer ER status using classical ML models.
  • RNA expression is the strongest single-omic predictor.
  • Multi-omic integration offers consistent, albeit modest, performance improvements.
  • Random Forest is a highly effective model for this type of high-dimensional biological data.

Original post by Priyanka Paudel, Madan Baduwal

"arXiv:2607.16250v1 Announce Type: new Abstract: Estrogen Receptor (ER) status is a critical biomarker in breast cancer diagnosis, prognosis, and treatment selection. Recent advances in high-throughput sequencing technologies have enabled the generation of multi-omics datasets tha…"

View on X

Originally posted by Priyanka Paudel, Madan Baduwal on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses