Related Experiment Video
Updated: Jan 2, 2026

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
Published on: March 1, 2024
A High-Performance Computing Implementation of Iterative Random Forest for the Creation of Predictive Expression
Ashley Cliff1,2, Jonathon Romero1,2, David Kainer2
1Bredesen Center for Interdisciplinary Research and Graduate Education, University of Tennessee Knoxville, Knoxville, TN 37996, USA.
We developed a high-performance computing implementation of Iterative Random Forest (iRF) for analyzing large biological datasets. This enables explainable AI eQTL analysis and creation of large Predictive Expression Networks.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Biological datasets are rapidly growing, necessitating advanced computational methods.
- Existing tools may not scale efficiently for large-scale genomic analyses.
Purpose of the Study:
- To present a high-performance computing (HPC)-capable implementation of Iterative Random Forest (iRF).
- To introduce a new method, iRF Leave One Out Prediction (iRF-LOOP), for creating large Predictive Expression Networks.
- To enable explainable AI eQTL analysis on massive SNP sets.
Main Methods:
- Developed an HPC-capable implementation of Iterative Random Forest (iRF).
- Introduced iRF Leave One Out Prediction (iRF-LOOP) for Predictive Expression Network construction.
- Benchmarked iRF performance against the R version on supercomputers Summit and Titan.
Main Results:
- The new iRF implementation significantly accelerates analysis of large SNP sets (over 1 million SNPs).
- iRF-LOOP successfully created Predictive Expression Networks for over 40,000 genes.
- Demonstrated iRF-LOOP's capability to identify biologically significant results.
Conclusions:
- The new HPC-capable iRF implementation allows for unprecedented scale in biological data analysis.
- iRF-LOOP provides a powerful tool for constructing complex Predictive Expression Networks.
- This advancement facilitates deeper insights into complex biological systems through large-scale genomic analysis.
More Related Videos
07:13Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
Published on: April 18, 2025
07:35Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Related Concept Videos
Randomized Experiments
Simple randomization
Simple...
Statistical Software for Data Analysis and Clinical Trials
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Survival Tree
Building a Survival Tree
Constructing a...
Random Sampling Method
Improving Translational Accuracy