Related Experiment Videos
Diagnostic accuracy of a DenseNet-121 deep learning algorithm for chest radiograph triage in health assessment
Lochan Shrestha1, Dinesh Maharjan2, Uttam Bista2
1Department of Radiology and Imaging, Patan Academy of Health Sciences, Lalitpur, Bagmati, Nepal lochanshrestha@pahs.edu.np.
Objectives:
To evaluate the diagnostic accuracy of a publicly available DenseNet-121 convolutional neural network (TorchXRayVision) for triaging chest radiographs of health assessment applicants at a tertiary hospital in Nepal.
Design:
Prospective, single-centre, shadow-mode diagnostic accuracy validation study. Reported in accordance with the Standards for Reporting of Diagnostic Accuracy Studies (STARD) 2015 checklist and STARD-Artificial Intelligence (AI)/Developmental and Exploratory Clinical Investigations of DEcision support systems driven by Artificial Intelligence (DECIDE-AI) guidelines.
Setting:
Department of Radiology and Imaging, Patan Academy of Health Sciences/Patan Hospital, Lalitpur, Nepal.
Participants:
826 consecutive health assessment applicants (foreign employment predeparture medical examination and student migration) undergoing chest radiography from 5 June 2026 to 20 June 2026 inclusive (16 days). Two cases were excluded due to Digital Imaging and Communications in Medicine technical failure.
Index Test:
DenseNet-121 algorithm (TorchXRayVision library, densenet121-res224-all pretrained weights). A maximum aggregated pathology probability score was derived per radiograph and compared against a post hoc derived threshold of 0.6258 (selected as the highest threshold achieving the prespecified ≥95% sensitivity criterion).
Reference Standard:
Single-reader-per-case review by one of three radiologists-two board-certified radiodiagnosticians (LS: 276 cases; DM: 275 cases) and one radiology resident (UB: 275 cases)-each blinded to AI output, using a standardised data collection worksheet capturing binary classification (abnormal/normal) and free-text findings.
Results:
Of 826 radiographs, 41 (4.97%) were classified as abnormal by the reference standard. At the post hoc derived threshold of 0.6258, the DenseNet-121 algorithm achieved: sensitivity 95.12% (95% CI 83.9% to 98.7%), specificity 77.2% (95% CI 74.1% to 80.0%), area under the receiver operating characteristic curve 0.9583 (95% bootstrap CI 0.9225 to 0.9843), negative predictive value (NPV) 99.67% (95% Wilson CI 98.8% to 99.9%), positive predictive value 17.89% (95% Wilson CI 13.4% to 23.5%) and Cohen's κ 0.237 (95% bootstrap CI 0.174 to 0.304). Brier score was 0.3621 (null Brier 0.0472) and expected calibration error was 0.564, confirming calibration failure due to score compression (range 0.52-0.72) despite preserved discrimination. The sensitivity estimate should be interpreted with caution given the relatively small number of reference-standard positives (n=41); the Wilson CI width of 14.8 percentage points (83.9-98.7%) reflects substantial uncertainty around this point estimate.
Cross-Validated Results:
10-fold cross-validation yielded bias-corrected sensitivity 95.12% (95% Wilson CI 83.9% to 98.7%; optimism 0.00 pp) and specificity 75.80% (95% Wilson CI 72.7% to 78.7%; optimism+1.40 pp), confirming by internal validation that primary metrics are not materially inflated by circular optimisation; independent external validation was not performed.
Conclusions:
The DenseNet-121 algorithm demonstrated high point-estimate sensitivity and excellent discrimination for chest radiograph triage in a Nepali health-assessment population, supporting its potential as a radiographic abnormality rule-out triage tool (NPV 99.67%); this does not constitute microbiological exclusion of active pulmonary tuberculosis. Systematic score compression-preserved discrimination despite calibration shift-is a quantifiable marker of low- and middle-income country distributional shift. Prospective local calibration studies and independent external validation are warranted before operational deployment.