Related Experiment Video
Updated: Jun 16, 2026

Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases
Published on: May 17, 2019
Omics-based Cancer Prognosis Across Ethnic Groups: From Feature Engineering to Disparity Detection and Mitigation
Teena Sharma1, Nishchal K Verma2, Yan Cui3
1The University of Tennessee Health Science Center, Memphis, TN 38105, USA and Indian Institute of Technology Guwahati, Assam 781039, India.
Abstract:
Artificial Intelligence (AI) models replicate human decision-making processes across numerous applications, including precision medicine using biomedical datasets. However, existing approaches suffer from two major challenges: unequal representation of ethnic groups resulting in data inequality, and the high dimensional nature of data, limiting generalization and interpretation. Such datasets hamper the performance of AI models for underrepresented or data disadvantaged groups, where the large number of features further complicates analysis. This paper introduces a machine learning approach for omics-based cancer prognosis across ethnic groups, employing an autoencoder based feature engineering method for feature extraction to detect disparities arising from biomedical data inequality and mitigate these disparities using transfer learning. The presented approach utilizes the feature engineering method to select and extract informative features, followed by machine learning schemes across ethnic groups, incorporating mixture learning, independent learning, and transfer learning that consider diverse group compositions to identify performance gaps and improve outcomes for underrepresented or data disadvantaged groups. Experimental analysis demonstrates that 22.34% of machine learning tasks exhibit performance disparities, which are mitigated using transfer learning for African American data disadvantaged group of mRNA Expression features from The Cancer Genome Atlas dataset. Similarly, across DNA methylation and MicroRNA Expression omics datasets, the presented approach effectively identified and reduced disparities in data disadvantaged groups compared to conventional feature engineering methods. Furthermore, Statistical tests validate the efficacy of the presented approach in detecting these disparities and enhancing the performance using transfer learning to mitigate the effects of these disparities across diverse groups.