Related Experiment Video
Updated: Aug 7, 2026

05:41
Synchronous Triplanar Reconstruction Integrated with Color Doppler Mapping for Precise and Rapid Localization of Thyroid Lesions
Published on: February 9, 2024
Image-based AI for Automated Diagnosis and Clinical Activity Grading in Thyroid Eye Disease: A Systematic Review and
Rui Hu1, Junzhe Zhao2, Guang-Yu Li1
1Department of Ophthalmology, the Second Hospital of Jilin University, Changchun 130000, China.
Ophthalmology
|August 5, 2026
Summary
Image-based AI shows promising accuracy for diagnosing thyroid eye disease (TED) and grading its activity. While a helpful diagnostic aid, it requires clinical supervision and further validation before autonomous use.
Area of Science:
- Ophthalmology
- Artificial Intelligence
- Medical Imaging
Background:
- Thyroid eye disease (TED) diagnosis and activity grading are challenging due to overlapping signs and subjective clinical activity scoring (CAS).
- The diagnostic accuracy of image-based artificial intelligence (AI) for TED has not been consolidated.
- AI presents a potential adjunct for improving diagnostic consistency and accuracy in TED management.
Purpose of the Study:
- To evaluate the diagnostic accuracy of image-based AI for TED diagnosis.
- To assess AI's accuracy in grading TED clinical activity (CAS ≥3 vs <3).
- To compare the performance of AI with ophthalmologists in TED diagnosis and activity grading.
Main Methods:
- A systematic literature search was conducted across PubMed, Embase, Web of Science, and Cochrane Library up to November 13, 2025.
- Included studies utilized image-based AI for TED diagnosis or activity grading.
- Bivariate random-effects models were used to pool sensitivity, specificity, and AUC, with risk of bias assessed using PROBAST+AI and evidence certainty using GRADE.
Main Results:
- 19 studies (n=8,744) were included, evaluating TED diagnosis (8 studies) and activity grading (11 studies).
- For TED diagnosis, AI achieved pooled sensitivity of 0.93, specificity of 0.84, and AUC of 0.93. For activity grading, AI achieved sensitivity of 0.82, specificity of 0.87, and AUC of 0.90.
- AI demonstrated higher sensitivity than ophthalmologist comparators for both diagnosis and activity grading, though comparator studies had limitations. Negative predictive values were high (LR- 0.09 for diagnosis, 0.21 for activity grading).
Conclusions:
- Image-based AI shows promising accuracy for TED diagnosis and CAS-based activity grading, potentially serving as a diagnostic rule-out aid under supervision.
- Current evidence has low certainty, substantial heterogeneity, and limitations in comparator studies, precluding autonomous AI deployment.
- CAS-based activity grading should not be used to rule out low-CAS active TED. Further prospective, multi-center, multi-ethnic validation is necessary.

