Related Experiment Videos
A robust machine learning approach for breast cancer subtype classification using relative gene expression order
Shamita Uma Kandan1, Osman Abul2
1Department of Electrical Engineering, College of Engineering, University of Sharjah, Sharjah, United Arab Emirates.
Computer Methods and Programs in Biomedicine
|June 16, 2026
Summary
Machine learning models using relative gene expression order effectively classify breast cancer subtypes. This approach enhances robustness and generalizability across diverse datasets, improving clinical decision-making.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Breast cancer subtype classification is crucial for personalized treatment and prognosis.
- Machine learning (ML) models analyzing gene expression show promise but struggle with technical data variability.
- Robust and generalizable ML classifiers are needed for reliable breast cancer subtyping.
Purpose of the Study:
- To develop and evaluate a robust machine learning approach for breast cancer subtype classification.
- To leverage relative gene expression order representations to overcome technical variability in gene expression data.
- To assess the performance and clinical relevance of these ML models across diverse datasets.
Main Methods:
- Proposed a machine learning approach using rank- and word2vec embedding-based relative gene expression order.
- Representations were derived from the within-sample relative expression order of PAM50 genes.
- Models were trained and systematically evaluated for robustness and clinical relevance on benchmark datasets (SCAN-B, TCGA-BRCA).
Main Results:
- Rank- and word2vec models achieved high precision and recall (≥91%) on cross-validation and test sets.
- Models demonstrated high accuracy (≥95%) on the external TCGA-BRCA dataset.
- Gene expression order representations preserved biological variation and enabled meaningful survival stratification.
Conclusions:
- Relative gene expression order, via ranks and word2vec embeddings, enables robust breast cancer subtype classification.
- This method reduces dependence on cohort-wide normalization, addressing technical variability.
- The approach supports reliable subtype classification in diverse clinical settings.