Related Experiment Video
Updated: Nov 7, 2025

Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases
Published on: May 17, 2019
Phenomapping of Patients with Primary Breast Cancer Using Machine Learning-Based Unsupervised Cluster Analysis
Sara Ferro1, Daniele Bottigliengo1, Dario Gregori1
1Unit of Biostatistics, Epidemiology and Public Health, Department of Cardiac Thoracic Vascular Sciences and Public Health, University of Padova, Via Loredan 18, 35121 Padova, Italy.
Unsupervised machine learning, specifically hierarchical agglomerative clustering, effectively identified two distinct patient subgroups in primary breast cancer (PBC). These clusters differ in age, hormone receptor status (estrogen and progesterone), and cathepsin D levels, aiding in understanding PBC heterogeneity.
Area of Science:
- Oncology
- Bioinformatics
- Computational Biology
Background:
- Primary breast cancer (PBC) presents significant heterogeneity across clinical, histopathological, and molecular features.
- Accurate classification of PBC is crucial for identifying patient subgroups and optimizing management strategies.
- Machine learning offers potential for deciphering complex relationships within heterogeneous clinical data.
Purpose of the Study:
- To explore the utility of unsupervised learning techniques for enhancing the classification of primary breast cancer.
- To identify distinct patient subgroups within PBC based on biological prognostic parameters.
Main Methods:
- Utilized a dataset comprising 712 women diagnosed with primary breast cancer.
- Applied four unsupervised clustering methods: K-means, self-organising maps, hierarchical agglomerative clustering (HAC), and Gaussian mixture models.
- Evaluated clustering performance, with HAC demonstrating superior results.
Main Results:
- Hierarchical agglomerative clustering identified two distinct patient clusters with differing clinical profiles.
- Cluster 1 comprised younger patients with lower estrogen receptor (ER) and progesterone receptor (PgR) values and lower cathepsin D levels.
- Cluster 2 included older patients with higher ER and PgR values. Age, ER, and PgR were identified as the most significant variables by HAC.
Conclusions:
- Unsupervised learning, particularly HAC, is a valuable tool for analyzing heterogeneous PBC data.
- This approach offers new perspectives for dissecting clinical heterogeneity in breast cancer research.
- The identified clusters highlight distinct patient profiles that may inform personalized treatment strategies.
More Related Videos
09:21Author Spotlight: Generating Neuronal Phenotypic Profiles - A Protocol to Culture and Image Human Midbrain Dopaminergic Neurons
Published on: July 7, 2023
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018