Related Experiment Video
Updated: Mar 10, 2026

Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases
Published on: May 17, 2019
Uncovering Clinically Relevant Breast Cancer Subtypes Biomarkers Using Integrative Bioinformatics and Machine
Prashansha Goel1, Nilofer Shaikh1
1Bioinformatics Department, Biotecnika Info Labs Pvt Ltd, Bengaluru, Karnataka, India.
Abstract:
A precise diagnosis and customized treatment become more difficult by the genomic heterogeneity of breast cancer (BRCA). In order to examine gene expression data from two separate Gene Expression Omnibus (GEO) microarray datasets, we used a integrative approach in this study that combined bioinformatics and machine learning. We were able to distinguish between universal and subtype-specific transcriptome patterns by identifying both common and subtype-specific differentially expressed genes (DEGs) using dual-level differential expression analysis. Functional enrichment analysis and the creation of protein-protein interaction networks identified important hub genes, including TPM3, MYLK, and COL17A1, which showed substantial dysregulation and were linked to high mutation rates and a bad prognosis. Survival analyses, which identified COL17A1 as a predictive predictor for the general population and MYLK for the Luminal B subtype, highlighted the clinical significance of these hub genes. We used both Random Forest and K-Nearest Neighbors classifiers to ensure robust biomarker identification. In the analysis, we prioritized 35 model-agnostic biomarkers that performed well in subtype categorization, such as PNMT and KRTAP10-8. This dual-model approach improved the reliability of biomarker identification while reducing model-specific biases. These results set the stage for early identification, more accurate subtype classification, and possible therapeutic targeting in breast cancer.

