Related Experiment Video
Updated: Feb 7, 2026

Deep Learning-Based Segmentation of Cryo-Electron Tomograms
Published on: November 11, 2022
Classification accuracy of a hierarchical molecular inference-based deep-learning system for CNS tumour diagnosis: a
H Lalchungnunga1, Christopher H Dampier1, Omkar Singh1
1Laboratory of Pathology, Center for Cancer Research, National Cancer Institute, National Institutes of Health, Bethesda, MD, USA.
Background:
Recent advances in artificial intelligence (AI) and computer vision empower deep-learning models to infer molecular features from histopathological images to classify CNS tumours. The aim of this study was to test the classification accuracy of a molecular inference-based AI assistant for CNS tumour diagnosis.
Methods:
In this multi-institutional, retrospective study, we used data from whole slide images of samples from patients aged 0-95 years, diagnosed with primary or recurrent CNS tumours. Reference diagnostic labels were determined by DNA methylation-based tumour classification to match one of 52 tumour types selected to encompass most types of gliomas, embryonal tumours, and meningeal and mesenchymal tumours encountered in clinical practice. The Neuropath-AI model was trained on 5835 samples from the National Cancer Institute (NCI; USA), the Children's Brain Tumor Network (USA), and the Digital Brain Tumour Atlas (Austria) to infer molecular features from whole slide images and to use these to predict tumour types with associated confidence scores. The test cohort comprised 5516 samples identified in laboratory archives between May 17, 2024, and May 13, 2025, from the NCI, Northwestern Medicine (USA), University of Pittsburgh Medical Center (USA), and University College London (UK). There were 2753 (50%) female and 2763 (50%) male patients, median age 43 years (IQR 25-59). The primary objective was to measure the classification accuracy of the model family-level and terminal classification predictions in test samples, with coprimary endpoints of sample coverage and prediction and balanced accuracy. Sample coverage was defined as samples receiving a model prediction with a confidence score above a specified threshold. Prediction accuracy and balanced accuracy were analysed in the covered samples (ie, those meeting the confidence criterion) and evaluated by comparing the top-1 or top-2 predictions with reference labels.
Findings:
Family-level classifications were reached in 5299 (96%) of 5516 samples. Predictions to one of the terminal classifications with at least moderate confidence were reached for 4772 (87%) samples. The single highest-scoring classification matched the reference label in 3817 (95% CI 3770-3865; 80% [95% CI 79-81]) of 4772 samples (balanced accuracy 66% [95% CI 63-70]). The two highest-scoring classifications contained the reference label in 4103 (95% CI 4056-4152; 86% [95% CI 85-87]) of 4772 samples (balanced accuracy 75% [95% CI 71-78]).
Interpretation:
Our model provides the basis for a clinically applicable deep-learning assistant to improve human efficiency and accuracy of CNS tumour diagnosis. The model will be made publicly available and could be implemented to assist human pathologists in future prospective studies.
Funding:
The Intramural Research Program of the National Institutes of Health.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Uncertainty in Measurement: Accuracy and Precision
Accuracy and Precision
Classification of Titrimetric Analysis Based on Reaction Types
Titrations between an acid and a base lead to neutralization reactions that form...
Cardiovascular Drugs: Classification based on Therapeutic Indications

