Related Experiment Video
Updated: Aug 2, 2026

Optimization of a Multiplex RNA-based Expression Assay Using Breast Cancer Archival Material
Published on: August 1, 2018
A robust machine learning approach for breast cancer subtype classification using relative gene expression order
Shamita Uma Kandan1, Osman Abul2
1Department of Electrical Engineering, College of Engineering, University of Sharjah, Sharjah, United Arab Emirates.
Background And Objective:
Breast cancer subtype classification is critical for clinical decision-making. It informs prognostic assessment and guides personalized treatment planning. Machine learning has been widely applied to develop subtype classifiers based on gene expression profiles due to its ability to capture complex molecular patterns. However, technical variation in gene expression data presents a major challenge in building classifiers that maintain robust and generalizable performance across different platforms and patient cohorts.
Methods:
We propose a robust machine learning approach leveraging relative gene expression order representations for subtype classification. Two representations: rank- and word2vec embedding-based, were evaluated to capture biological variation across subtypes with minimal cross-sample dependence. These representations were derived from the within-sample relative expression order of PAM50 genes, a well-established 50-gene signature for defining breast cancer subtypes. Machine learning models were trained on these representations and systematically evaluated for robustness and clinical relevance across benchmark datasets.
Results:
Both rank- and word2vec embedding-based machine learning models demonstrated robust performance during cross-validation on the SCAN-B training set and on SCAN-B internal and semi-external test sets, achieving at least 91% precision and recall. On the fully external TCGA-BRCA dataset, the models accurately classified the majority of subtypes, with at least 95% accuracy. Predicted luminal subtypes also showed clinically meaningful stratification in survival analyses, and both representations effectively preserved biological variation, as reflected in subtype-wise sample segregation in PCA visualizations. In addition, word2vec-derived gene embeddings captured biologically meaningful co-occurring gene clusters.
Conclusion:
Ranks and word2vec embeddings derived from relative gene expression order enable machine learning models to robustly classify breast cancer subtypes. They address technical variability by reducing dependence on cohort-wide normalization, thereby supporting reliable subtype classification in diverse clinical settings.

