Related Experiment Video
Updated: Sep 13, 2025

Constructing and Visualizing Models using Mime-based Machine-learning Framework
Published on: July 22, 2025
Development and validation of machine learning-based diagnostic models using blood transcriptomics for early
Xin Huang1,2, Di Ouyang3, Weiming Xie4
1The First Clinical Medical College, Nanjing University of Chinese Medicine, Nanjing, China.
Insights
Early prediction of Type 1 Diabetes Mellitus (T1DM) in children is possible using blood transcriptomics and machine learning. This approach identifies predictive signatures up to 46 months before diagnosis, paving the way for non-invasive diagnostic tools.
Area of Science:
- Genomics and Bioinformatics
- Machine Learning in Healthcare
- Pediatric Endocrinology
Background:
- Early identification of Type 1 Diabetes Mellitus (T1DM) in children is critical for timely interventions.
- Peripheral blood transcriptomic analysis offers a minimally invasive method for early biomarker discovery.
- Predicting T1DM onset in children up to 46 months before clinical diagnosis is the focus.
Purpose of the Study:
- To develop and validate machine learning algorithms for predicting T1DM onset in children.
- To utilize transcriptomic signatures from peripheral blood for early T1DM prediction.
- To identify predictive biomarkers for T1DM in pediatric populations.
Main Methods:
- Analyzed RNA-sequencing data from 247 pre-diabetic children and healthy controls.
- Employed five feature selection methods and nine machine learning algorithms to create 45 model combinations.
- Validated models using quantitative polymerase chain reaction (qPCR) in an independent cohort.
Main Results:
- Significant differential gene expression patterns were observed between pre-diabetic and control groups.
- Four model combinations showed superior predictive performance, accurately predicting T1DM onset up to 46 months prior.
- Elastic Net-based models achieved perfect classification in the validation cohort, indicating clinical viability.
Conclusions:
- Peripheral blood transcriptomics combined with machine learning enables early pediatric T1DM prediction.
- Identified transcriptomic signatures and validated models form a basis for non-invasive diagnostic tools.
- Findings support precision medicine for childhood diabetes prevention, requiring larger cohort validation.
Background:
Early identification of Type 1 Diabetes Mellitus (T1DM) in pediatric populations is crucial for implementing timely interventions and improving long-term outcomes. Peripheral blood transcriptomic analysis provides a minimally invasive approach for identifying predictive biomarkers prior to clinical manifestation. This study aimed to develop and validate machine learning algorithms utilizing transcriptomic signatures to predict T1DM onset in children up to 46 months before clinical diagnosis.
Methods:
We analyzed 247 peripheral blood RNA-sequencing samples from pre-diabetic children and age-matched healthy controls. Differential gene expression analysis was performed using established bioinformatics pipelines to identify significantly dysregulated transcripts. Five feature selection methods (Lasso, Elastic Net, Random Forest, Support Vector Machine, and Gradient Boosting Machine) were employed to optimize gene sets. Nine machine learning algorithms (Decision Tree, Gradient Boosting Machine, K-Nearest Neighbors, Linear Discriminant Analysis, Logistic Regression, Multilayer Perceptron, Naive Bayes, Random Forest, and Support Vector Machine) were combined with selected features, generating 45 unique model combinations. Performance was evaluated using accuracy, precision, recall, and F1-score metrics. Model validation was conducted using quantitative polymerase chain reaction (qPCR) in an independent cohort of six children (three healthy, three diabetic).
Results:
Transcriptomic analysis revealed significant differential expression patterns between pre-diabetic and control groups. Four model combinations demonstrated superior predictive performance: Lasso+K-Nearest Neighbors, Elastic Net + K-Nearest Neighbors, Elastic Net + Random Forest, and Support Vector Machine+K-Nearest Neighbors. These models achieved high accuracy in predicting diabetes onset up to 46 months before clinical diagnosis. Both Elastic Net-based models achieved perfect classification performance in the validation cohort, demonstrating their potential as clinically viable diagnostic tools.
Conclusion:
This study establishes the feasibility of integrating peripheral blood transcriptomic profiling with machine learning for early pediatric T1DM prediction. The identified transcriptomic signatures and validated predictive models provide a foundation for developing clinically translatable, non-invasive diagnostic tools. These findings support the implementation of precision medicine approaches for childhood diabetes prevention and warrant validation in larger, multi-center cohorts to assess generalizability and clinical utility.

