Related Experiment Video
Updated: Aug 27, 2025

07:42
A Data-Driven Approach to Quantifying Immune States in Sepsis
Published on: February 7, 2025
302
An application based on bioinformatics and machine learning for risk prediction of sepsis at first clinical
Songchang Shi1, Xiaobin Pan1, Lihui Zhang1
1Department of Critical Care Medicine, Shengli Clinical Medical College of Fujian Medical University, Fujian Provincial Hospital South Branch, Fujian Provincial Jinshan Hospital, Fujian Provincial Hospital, Fuzhou, China.
Frontiers in Genetics
|September 26, 2022
Summary
This study introduces a machine learning workflow using transcriptomic data to predict sepsis risk. The CatBoost model, combined with SHAP analysis, effectively identifies key genes for early sepsis detection.
Area of Science:
- Bioinformatics and Computational Biology
- Genomics and Transcriptomics
- Machine Learning in Medicine
Background:
- Predicting disease risk from genotypic changes to phenotypic traits using machine learning presents significant challenges.
- Transcriptomic data offers potential for early disease risk prediction but requires sophisticated analytical approaches.
- Existing methods face difficulties in accurately linking molecular data to clinical outcomes.
Purpose of the Study:
- To develop and validate a bioinformatics and machine learning workflow for predicting sepsis risk using transcriptomic data.
- To overcome challenges in disease risk prediction by integrating transcriptomic data with advanced computational methods.
- To identify optimal genes and build a robust model for early sepsis risk assessment.
Main Methods:
- Processing and annotating high-throughput sequencing transcriptomic data using R software.
- Constructing and evaluating machine learning models in Python, including feature selection via recursive elimination.
- Utilizing Shapley Additive explanation (SHAP) for model interpretation and visualization of gene significance.
Main Results:
- Identification of the top 10 optimal genes for sepsis risk prediction using machine learning-based feature selection.
- Selection of the CatBoost model as the optimal performer among evaluated machine learning models.
- Detailed exploration of individual gene significance and inter-gene interactions within the predictive model via SHAP analysis.
Conclusions:
- The combined use of CatBoost and SHAP provides a high-performing machine learning model for predicting sepsis risk from transcriptomic data.
- The developed workflow offers a novel approach for investigating gene-disease mechanisms and improving sepsis risk prediction.
- This study highlights the potential of integrating bioinformatics and machine learning for advancing precision medicine in sepsis management.

