Related Experiment Video
Updated: Dec 18, 2025

Cancer-Associated Fibroblasts from Mouse Mammary Tumors as Tools for Molecular and Computational Studies
Published on: July 3, 2025
Assessing reproducibility and veracity across machine learning techniques in biomedicine: A case study using TCGA
Ahyoung Amy Kim1, Samir Rachid Zaim2, Vignesh Subbian3
1Graduate Interdisciplinary Program in Statistics and Data Science, The University of Arizona, United States.
Identifying gene biomarkers for drug development faces challenges due to poor reproducibility in statistical learning methods. This study highlights the need for transparent analysis to improve clinical validity of gene expression data.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Gene biomarker identification for drug development often lacks clinical validity and reproducibility.
- Statistical learning tools are crucial for genomic data analysis, necessitating an examination of their limitations.
Purpose of the Study:
- To demonstrate methodological gaps in common statistical learning techniques for gene expression analysis.
- To assess the classification ability and reproducibility of machine learning models for gene biomarker detection.
Main Methods:
- Trained six machine learning models on The Cancer Genome Atlas (TCGA) cancer data.
- Evaluated classification ability using standard performance metrics (specificity, sensitivity, precision, F1 score).
- Assessed reproducibility by quantifying the consistency of selected gene classifiers.
Main Results:
- Random forest models exhibited the best overall classification ability.
- Few genes were consistently selected across different methods, indicating poor identifiability and reproducibility.
- Inherent differences in statistical machine learning methods challenge the reproducibility of gene expression discoveries.
Conclusions:
- Statistical machine learning models show significant variation in high-dimensional gene expression data.
- Transparent analysis procedures, including preprocessing, parameterization, and model selection, are essential for clinical validity and utility.
More Related Videos
07:47Author Spotlight: Unveiling Transmembrane Protein Family-Related Markers in Gastric Cancer and Implications for Targeted Therapies
Published on: September 15, 2023
07:41Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases
Published on: May 17, 2019