Related Experiment Video
Updated: Nov 4, 2025

Constructing and Visualizing Models using Mime-based Machine-learning Framework
Published on: July 22, 2025
Permutation-based identification of important biomarkers for complex diseases via machine learning models
Xinlei Mi1, Baiming Zou2, Fei Zou2
1Department of Preventive Medicine, Northwestern University, Feinberg School of Medicine, Chicago, IL, USA.
We introduce Permutation-based Feature Importance Test (PermFIT) to interpret complex machine learning models for human disease studies. PermFIT identifies key biomarkers, enhancing disease prevention, diagnosis, and treatment strategies.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Human disease research faces challenges due to complex molecular mechanisms and genetic factors.
- Machine learning models offer powerful analytical tools but often lack transparency, hindering biomarker identification.
- Interpreting individual features in complex models is crucial for advancing disease understanding and therapeutic strategies.
Purpose of the Study:
- To develop and validate a computationally efficient method, Permutation-based Feature Importance Test (PermFIT), for estimating and testing feature importance in complex machine learning frameworks.
- To enhance the interpretability of sophisticated algorithms like deep neural networks, random forests, and support vector machines.
- To aid researchers in identifying critical biomarkers for complex human diseases.
Main Methods:
- Developed Permutation-based Feature Importance Test (PermFIT) for feature importance estimation and testing.
- Implemented PermFIT in a computationally efficient manner, avoiding model refitting.
- Applied PermFIT to various machine learning models, including deep neural networks, random forests, and support vector machines.
Main Results:
- PermFIT provides valid statistical inference for feature importance.
- The method improves the prediction accuracy of machine learning models.
- Extensive numerical studies demonstrated PermFIT's effectiveness across diverse scenarios.
Conclusions:
- PermFIT offers a practical and efficient approach to interpreting complex machine learning models in biomedical research.
- The tool successfully identified important biomarkers in real-world datasets, such as the Cancer Genome Atlas kidney tumor data and the HITChip atlas data.
- PermFIT aids in hypothesis generation for disease prevention, diagnosis, and treatment by enhancing biomarker discovery and model performance.
More Related Videos
07:35Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
08:51Author Spotlight: Integrated Multi-Omics Analysis for Unveiling Multicellular Immune Signatures in Clinical Heart Attack Cohorts
Published on: September 20, 2024