Related Experiment Video
Updated: Apr 12, 2026

14:27
Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data
Published on: June 26, 2013
15.6K
High-dimensional Biomarker Identification for Scalable and Interpretable Disease Prediction via Machine Learning
Yifan Dai1, Fei Zou1,2, Baiming Zou1,3
1Department of Biostatistics, University of North Carolina at Chapel Hill, Chapel Hill, NC, USA.
Biorxiv : the Preprint Server for Biology
|October 17, 2024
Summary
This study introduces the High-dimensional Feature Importance Test (HiFIT) to identify key biomarkers and clinical risk factors from complex omics and clinical data for disease research. HiFIT improves disease prediction and biomarker discovery, aiding precision medicine.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Complex human diseases are influenced by both omics data and clinical features.
- Identifying key biomarkers and risk factors is crucial for understanding disease mechanisms and advancing precision medicine.
- High-dimensionality and complex associations in omics data pose analytical challenges.
Purpose of the Study:
- To develop a robust framework for biomarker and clinical risk factor identification from high-dimensional omics and clinical data.
- To enhance disease outcome prediction and feature importance identification for complex diseases.
- To provide a computationally efficient tool for researchers in precision medicine.
Main Methods:
- Proposed an ensemble data-driven biomarker identification tool, Hybrid Feature Screening (HFS), to construct a candidate feature set.
- Developed the High-dimensional Feature Importance Test (HiFIT) framework, refining candidate features using a permutation-based importance test.
- Utilized extensive numerical simulations and real-world applications to validate the framework.
Main Results:
- HiFIT demonstrated superior performance in outcome prediction compared to existing methods.
- The framework effectively identified key biomarkers and clinical risk factors.
- Demonstrated computational efficiency in handling high-dimensional data.
Conclusions:
- The HiFIT framework offers a powerful approach for biomarker discovery and risk factor identification in complex diseases.
- HiFIT facilitates advancements in early disease diagnosis and personalized treatment strategies.
- An R package for HiFIT is publicly available to support the research community.

