Related Experiment Video
Updated: Aug 7, 2026

Synchronous Triplanar Reconstruction Integrated with Color Doppler Mapping for Precise and Rapid Localization of Thyroid Lesions
Published on: February 9, 2024
Image-Based Artificial Intelligence for Automated Diagnosis and Clinical Activity Grading in Thyroid Eye Disease: A
Rui Hu1, Junzhe Zhao2, Guang-Yu Li1
1Department of Ophthalmology, The Second Hospital of Jilin University, Changchun, China.
Topic:
To evaluate image-based artificial intelligence (AI) accuracy for thyroid eye disease (TED) diagnosis and clinical activity grading (Clinical Activity Score [CAS] ≥ 3 vs. <3) and compare AI with ophthalmologists.
Clinical Relevance:
TED diagnosis and activity grading are challenging because of overlapping signs and subjective CAS (interrater variability up to 30%). Its consolidated accuracy is unquantified.
Methods:
We searched PubMed, Embase, Web of Science, and Cochrane Library through November 13, 2025, for studies of TED diagnosis or activity grading using image-based AI. Risk of bias was assessed with PROBAST+AI and certainty with Grading of Recommendations Assessment, Development, and Evaluation (GRADE). Bivariate random-effects models pooled sensitivity, specificity, and area under the curve (AUC).
Results:
Nineteen studies (n = 8744) were included: 8 evaluated diagnosis (11 validation sets; n = 4662) and 11 activity grading (11 validation sets; n = 4082). For diagnosis, sensitivity was 0.93 (95% confidence interval [CI], 0.86-0.96), specificity 0.84 (95% CI, 0.78-0.88), and AUC 0.93 (95% CI, 0.90-0.95); after outlier exclusion, 0.90, 0.82, and 0.92, respectively. For activity grading, sensitivity was 0.82 (95% CI, 0.77-0.85), specificity 0.87 (95% CI, 0.82-0.91), and AUC 0.90 (95% CI, 0.88-0.91); after outlier exclusion, 0.80, 0.84, and 0.89, respectively. AI showed higher sensitivity than available ophthalmologist comparators for diagnosis (P = 0.01) and activity grading (P = 0.003), but comparisons involved selected readers and limited clinical context. Negative likelihood ratio was 0.09 for diagnosis and 0.21 for activity grading. PROBAST+AI rated 12.5% of diagnosis studies and 36.4% to 54.5% of activity-grading studies at high risk of bias; GRADE certainty was low.
Conclusion:
Image-based AI showed promising accuracy for TED diagnosis and CAS-based activity grading. Evidence supports cautious use as a diagnosis rule-out aid under clinical supervision when negative results are interpreted with symptoms and clinical context; however, CAS-based activity grading should not rule out low-CAS active TED. Low certainty, substantial heterogeneity, and limitations of image-based comparisons preclude immediate autonomous clinical deployment. Patients with suspected TED should still be referred for ophthalmologic assessment. Prospective, multicenter, multiethnic validation is required.
Financial Disclosure(S):
The author has no/the authors have no proprietary or commercial interest in any materials discussed in this article.

