Related Experiment Video
Updated: Feb 20, 2026

Building Up a High-throughput Screening Platform to Assess the Heterogeneity of HER2 Gene Amplification in Breast Cancers
Published on: December 5, 2017
Machine learning approaches to decipher hormone and HER2 receptor status phenotypes in breast cancer
Emmanuel S Adabor1, George K Acquaah-Mensah2
1Stellenbosch University.
Abstract:
Breast cancer prognosis and administration of therapies are aided by knowledge of hormonal and HER2 receptor status. Breast cancer lacking estrogen receptors, progesterone receptors and HER2 receptors are difficult to treat. Regarding large data repositories such as The Cancer Genome Atlas, available wet-lab methods for establishing the presence of these receptors do not always conclusively cover all available samples. To this end, we introduce median-supplement methods to identify hormonal and HER2 receptor status phenotypes of breast cancer patients using gene expression profiles. In these approaches, supplementary instances based on median patient gene expression are introduced to balance a training set from which we build simple models to identify the receptor expression status of patients. In addition, for the purpose of benchmarking, we examine major machine learning approaches that are also applicable to the problem of finding receptor status in breast cancer. We show that our methods are robust and have high sensitivity with extremely low false-positive rates compared with the well-established methods. A successful application of these methods will permit the simultaneous study of large collections of samples of breast cancer patients as well as save time and cost while standardizing interpretation of outcomes of such studies.
Insights
New methods accurately predict breast cancer receptor status using gene expression profiles. This approach aids in treating difficult breast cancers and analyzing large patient data sets efficiently.
Area of Science:
- Oncology
- Bioinformatics
- Genomics
Background:
- Accurate determination of estrogen receptor (ER), progesterone receptor (PR), and HER2 receptor status is crucial for breast cancer prognosis and treatment selection.
- Breast cancers lacking these receptors are often more challenging to treat.
- Current wet-lab methods may not cover all samples in large genomic datasets like The Cancer Genome Atlas (TCGA).
Purpose of the Study:
- To develop and validate novel computational methods for identifying ER, PR, and HER2 receptor status phenotypes in breast cancer patients using gene expression profiles.
- To provide a robust and cost-effective alternative to traditional wet-lab methods for receptor status determination.
- To benchmark the performance of these new methods against established machine learning approaches.
Main Methods:
- Introduction of median-supplement methods utilizing patient gene expression profiles.
- Balancing training datasets by introducing supplementary instances based on median patient gene expression.
- Building simple predictive models to determine receptor expression status.
- Benchmarking against major machine learning approaches for receptor status identification.
Main Results:
- The proposed median-supplement methods demonstrate robustness and high sensitivity in predicting breast cancer receptor status.
- These methods achieve extremely low false-positive rates compared to well-established techniques.
- The computational approach offers a viable alternative for analyzing large-scale patient sample collections.
Conclusions:
- Median-supplement methods provide an accurate and efficient way to determine hormonal and HER2 receptor status in breast cancer patients from gene expression data.
- This approach can significantly save time and costs associated with traditional methods.
- Successful implementation allows for standardized interpretation and simultaneous study of extensive breast cancer patient cohorts.

