Related Experiment Video
Updated: Aug 7, 2026

Synchronous Triplanar Reconstruction Integrated with Color Doppler Mapping for Precise and Rapid Localization of Thyroid Lesions
Published on: February 9, 2024
Image-based AI for Automated Diagnosis and Clinical Activity Grading in Thyroid Eye Disease: A Systematic Review and
Rui Hu1, Junzhe Zhao2, Guang-Yu Li1
1Department of Ophthalmology, the Second Hospital of Jilin University, Changchun 130000, China.
Topic:
To evaluate diagnostic accuracy of image-based AI for thyroid eye disease (TED) diagnosis and clinical activity grading (CAS ≥3 vs <3), and to compare AI performance with ophthalmologists.
Clinical Relevance:
Diagnosis of TED and activity grading remain challenging due to overlapping signs and subjective CAS (inter-rater variability up to 30%). Image-based AI is a potential adjunct, but its consolidated accuracy is unquantified.
Methods:
We searched PubMed, Embase, Web of Science, and Cochrane Library to 13 November 2025 for studies using image-based AI for TED diagnosis or activity grading. Risk of bias was assessed with PROBAST+AI, evidence certainty with GRADE. Bivariate random-effects models pooled sensitivity, specificity, and AUC.
Results:
19 studies (n=8,744) were included: 8 evaluated diagnosis (11 validation sets; n=4,662) and 11 addressed activity grading (11 validation sets; n=4,082). For TED diagnosis, AI achieved sensitivity 0.93 (95% CI: 0.86-0.96), specificity 0.84 (95% CI: 0.78-0.88), and AUC 0.93 (95% CI: 0.90-0.95); after outlier exclusion, sensitivity was 0.90, specificity 0.82, and AUC 0.92. For activity grading, sensitivity was 0.82 (95% CI: 0.77-0.85), specificity 0.87 (95% CI: 0.82-0.91), and AUC 0.90 (95% CI: 0.88-0.91); after outlier exclusion, sensitivity was 0.80, specificity 0.84, and AUC 0.89. AI showed higher sensitivity than available ophthalmologist comparators for diagnosis (P = 0.01) and activity grading (P ≤ 0.003), but comparator experiments used selected reader groups, limited clinical context, and did not report annual TED case volume. LR- was 0.09 for diagnosis and 0.21 for activity grading. PROBAST+AI rated 12.5% of diagnosis studies and 36.4%-54.5% of activity-grading studies at high risk of bias; GRADE certainty was low.
Conclusion:
Image-based AI showed promising accuracy for TED diagnosis and CAS-based activity grading. Evidence supports its cautious use as a diagnosis rule-out aid under clinical supervision when a negative result is interpreted with symptoms and clinical context; however, CAS-based activity grading should not be used to rule out low-CAS active TED. Low certainty of evidence, substantial heterogeneity, and limitations of image-based comparator experiments preclude immediate autonomous clinical deployment. Patients with visual or periocular symptoms suspicious for TED should still be referred for ophthalmologic assessment. Prospective, multi-center, multi-ethnic validation is required.

