Related Experiment Video
Updated: Jun 16, 2026

07:41
Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases
Published on: May 17, 2019
Omics-based Cancer Prognosis Across Ethnic Groups: From Feature Engineering to Disparity Detection and Mitigation.
Teena Sharma1, Nishchal K Verma2, Yan Cui3
1The University of Tennessee Health Science Center, Memphis, TN 38105, USA and Indian Institute of Technology Guwahati, Assam 781039, India.
Summary
This study introduces an AI approach to address data inequality in cancer prognosis, using feature engineering and transfer learning to improve outcomes for underrepresented ethnic groups.
Area of Science:
- Biomedical data science
- Machine learning in healthcare
- Computational biology
Background:
- Artificial Intelligence (AI) models face challenges with unequal ethnic representation in biomedical data, leading to data inequality and poor generalization.
- High-dimensional omics data further complicates AI model analysis, particularly for underrepresented or data-disadvantaged groups.
Purpose of the Study:
- To develop and validate a machine learning approach for omics-based cancer prognosis that addresses data inequality across ethnic groups.
- To detect and mitigate performance disparities in AI models caused by imbalanced biomedical datasets.
- To improve AI model outcomes for underrepresented populations in precision medicine.
Main Methods:
- An autoencoder-based feature engineering method was employed for feature extraction from omics data.
- Machine learning schemes including mixture learning, independent learning, and transfer learning were applied across ethnic groups.
- Transfer learning was utilized to mitigate identified performance gaps for data-disadvantaged groups.
Main Results:
- Experimental analysis revealed performance disparities in 22.34% of machine learning tasks.
- Transfer learning effectively mitigated disparities for the African American group using mRNA Expression features from The Cancer Genome Atlas dataset.
- The approach successfully identified and reduced disparities in DNA methylation and MicroRNA Expression omics datasets for data-disadvantaged groups.
Conclusions:
- The presented approach demonstrates efficacy in detecting performance disparities in AI models due to biomedical data inequality.
- Transfer learning significantly enhances AI model performance and mitigates the effects of data disparities across diverse ethnic groups.
- This method offers a pathway to more equitable and accurate AI-driven cancer prognosis.