Related Experiment Video
Updated: Mar 2, 2026

Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases
Published on: May 17, 2019
Integrating Genomic and Nongenomic Data to Stratify the Risk of Contralateral Breast Cancer After Radiation Therapy
Sangkyu Lee1, Xiang Shu2, Andriy Derkach2
1Department of Radiation Oncology, NYU Grossman School of Medicine, New York, New York.
Purpose:
Women treated with radiation therapy (RT) for breast cancer have an increased risk of developing radiation-associated contralateral breast cancer (CBC). Predicting CBC events is challenging because of the complex interplay of genomic, treatment, personal, and clinical factors. This study investigated computational methods that integrate genome-wide single-nucleotide polymorphisms and nongenomic data to develop a risk stratification model for developing CBC in women treated with RT for their first primary breast cancer.
Methods And Materials:
This study used a subset of the population-based Women's Environmental Cancer and Radiation Epidemiology study that included 633 CBC cases and 1253 individually matched unilateral breast cancer controls who were treated with RT and had single-nucleotide polymorphism data available from a genome-wide association study. The study population was split into training, validation, and test sets for rigorous modeling and validation. Three data integration methods were compared in terms of their ability to stratify CBC risk: (1) naive integration; (2) sequential integration; and (3) sequential iterative integration. A biological analysis of the final model was performed using gene set enrichment analysis and protein-protein interaction analysis with gene annotation information informed by the model.
Results:
The best-performing integration method was the sequential iterative integration equipped with the mixed-effect random forest algorithm. This approach achieved an area under the curve of 0.64 to stratify CBC risk in the test set, representing moderate predictive power. Calibration analysis showed good agreement between the lowest and highest risk bins stratified using sorted predicted values in the test set, resulting in an odds ratio of 3.27 for both predicted and observed CBC occurrence. Gene set enrichment analysis and protein-protein interaction analysis revealed that genes with high importance scores were associated with pathways relevant to lipid and fatty acid metabolism as well as breast cancer sensitivity to tamoxifen.
Conclusions:
The mixed-effect random forest approach demonstrated the potential for integrating high-dimensional genomic and low-dimensional nongenomic data to stratify CBC risk.

