Predicting isocitrate dehydrogenase mutation status in acute myeloid leukemia from gene expression profiles by
Ina Jung1, Anne-Laure Vitte1, Florent Chuffart1
1Université Grenoble Alpes, INSERM U1209, CNRS UMR 5309, Institute for Advanced Biosciences, 38000, Grenoble, France.
None:
We developed machine-learning models to predict isocitrate dehydrogenase (IDH) mutation status in acute myeloid leukemia (AML) from gene expression profiles and to reconstruct missing IDH annotations across public datasets. Transcriptomic data from 19 cohorts (5844 samples) were harmonized using batch correction, and 1546 samples with known IDH status were used to train a feed-forward neural network and a logistic regression (LR) classifier within a nested cross-validation framework, followed by independent validation in the TCGA-LAML dataset. The LR model showed superior performance, achieving receiver operating characteristic area under the curve [Formula: see text] 0.994 [Formula: see text] 0.007, accuracy [Formula: see text] 0.983 [Formula: see text] 0.006, balanced accuracy [Formula: see text] 0.979 [Formula: see text] 0.005, sensitivity for the IDH-mutant (IDH-MUT) class [Formula: see text] 0.972 [Formula: see text] 0.010, and specificity [Formula: see text] 0.986 [Formula: see text] 0.008, and correctly classified all IDH-MUT cases in the independent cohort. Applying the final model to samples lacking annotations enabled reconstruction of IDH status for 4148 AML cases, expanding the number of molecularly characterized transcriptomes available for downstream analyses. Predicted groups recapitulated known IDH-associated transcriptional signatures, supporting biological validity. This work demonstrates that IDH mutation status can be accurately inferred from transcriptomic data alone and provides a scalable framework to recover missing genomic annotations, thereby enhancing the utility of public AML resources for large-scale biological and translational research.

