Related Experiment Video
Updated: Aug 13, 2026

A 3D Digital Model for the Diagnosis and Treatment of Pulmonary Nodules
Published on: May 19, 2023
Diagnostic performance of deep learning models in differentiating benign and malignant pulmonary nodules: a
Fei Chen1, Lujiao Chen2, Yurui Lv1
1School of Medicine, Shaoxing University, Shaoxing, China.
Background:
Lung cancer is the most lethal malignant tumor globally. Accurate differentiation between benign and malignant pulmonary nodules (PNs) is the core of early screening and diagnosis for lung cancer. Deep learning (DL) models can autonomously extract high-dimensional information from medical images, demonstrating technical advantages in this task. Our study systematically evaluates the accuracy and efficacy of DL models in differentiating benign from malignant PNs, providing systematic, high-quality integrated evidence for their diagnostic performance.
Methods:
Our study was conducted following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses of Diagnostic Test Accuracy Studies (PRISMA-DTA) reporting checklist. We searched the databases of PubMed, Cochrane Library, Embase, and Web of Science for relevant studies published up to 24 September 2025. Two researchers independently screened the literature according to inclusion and exclusion criteria, using the Quality Assessment of Diagnostic Accuracy Studies-2 (QUADAS-2) to assess the risk of bias and applicability of the included studies. The pooled sensitivity, specificity, and area under the curve (AUC) were used to evaluate the diagnostic performance of DL models. Subgroup analysis and meta-regression were employed to investigate sources of heterogeneity. Sensitivity analysis and publication bias tests were conducted concurrently, and a Fagan plot was generated to assess clinical utility.
Results:
A total of 21 retrospective studies were ultimately included. The QUADAS-2 quality assessment indicated that the overall risk of bias in the included studies was manageable. The combined sensitivity, specificity, and AUC of DL models for distinguishing between benign and malignant PNs were 0.92 [95% confidence interval (CI): 0.90-0.93], 0.90 (95% CI: 0.86-0.92), and 0.96 (95% CI: 0.94-0.98). The pooled sensitivity and specificity were heterogeneous (I2>75%). A threshold effect test confirmed that no significant threshold effect was observed in our study (Spearman's r=0.221, P=0.336). Subgroup analysis and meta-regression confirmed that model validation method was the primary source of heterogeneity in our study (P=0.00), with cross-validation models demonstrating significantly higher specificity than train-test split models (0.92 vs. 0.84, P=0.00). Models based on the Transformer architecture outperformed convolutional neural networks (CNNs) in terms of both sensitivity (0.93 vs. 0.92, P=0.00) and specificity (0.91 vs. 0.89, P=0.00). Sensitivity analysis confirmed the stability of the pooled results, and the funnel plot revealed no significant publication bias (P=0.56). The Fagan plot indicated that at a pretest probability of 50%, the model achieved a positive predictive post-test probability of 90% and a negative predictive post-test probability of only 8%, demonstrating good clinical utility.
Conclusions:
DL models demonstrate both high sensitivity and specificity in distinguishing benign from malignant PNs, exhibiting outstanding overall diagnostic performance. With prospective validation and technical optimization in the future, their clinical utility is expected to improve further.
