Related Experiment Video
Updated: Jul 1, 2026

10:06
Cell Type-specific Gene Expression Profiling in the Mouse Liver
Published on: September 17, 2019
Explainable machine learning identifies small high-performance gene signatures from full transcriptomes in human
Joseph Mugaanyi1,2, Jing Huang3, Sheng Ye2
1Health Science Center, Ningbo University, Ningbo, Zhejiang, China.
Scientific Reports
|June 29, 2026
Summary
Researchers compressed large liver transplant gene lists into compact panels for diagnostics. These panels show high accuracy within specific patient groups but lack broad applicability across different studies.
Area of Science:
- Genomics and Bioinformatics
- Transplant Medicine
- Machine Learning in Healthcare
Background:
- Transcriptomic studies in liver preservation and ischemia-reperfusion injury (IRI) generate extensive gene lists, hindering clinical translation.
- Genome-wide profiling is impractical for time-sensitive organ procurement decisions.
- A need exists for compact, interpretable gene panels deployable on targeted platforms like NanoString or RT-qPCR.
Purpose of the Study:
- To determine if human liver transcriptomes can be compressed into diagnostic gene panels without sacrificing classification performance.
- To evaluate the diagnostic accuracy and cross-dataset transferability of machine learning-derived gene panels.
Main Methods:
- Seven human liver transcriptomic datasets (GEO) were re-processed for machine learning (ML) analysis.
- Harmonized gene symbols (HGNC) were used for all modeling.
- Elastic Net logistic regression and XGBoost models were trained using nested cross-validation, with feature selection on the outer training partition. The union of top genes from both models formed the panel.
Main Results:
- Within-cohort classification achieved high Area Under the Curve (AUC) values (0.842-0.965) for datasets with sufficient sample size (n≥30).
- Specific cohorts (GSE151648_TIME, GSE12720) demonstrated excellent diagnostic performance (AUC=0.959-0.965) with high specificity and positive predictive value (>0.90).
- Leave-One-Dataset-Out validation revealed poor cross-dataset transferability (balanced accuracy=0.500, macro F1≤0.333), indicating context-specific panels.
- A subset of reperfusion-focused cohorts showed high transferability (AUC=0.99-1.00) with convergent gene signatures (inflammatory, proteotoxic stress, AP-1 axis).
Conclusions:
- Explainable ML models can effectively compress liver transplant transcriptomes into compact, context-specific gene panels with strong within-cohort diagnostic accuracy.
- Convergent gene signatures in reperfusion cohorts provide biological validation for these panels.
- The poor cross-dataset transferability highlights the necessity for cohort-tailored assay development and prospective validation.
