Related Experiment Videos
An in-silico comparative cross-sectional diagnostic- accuracy study evaluating artificial intelligence platforms in
Suhail H Serbaya1, Saud Hasan Surbaya2, Tamara Abdulrahman Hafiz3
1Department of Industrial Engineering, Faculty of Engineering, King Abdulaziz University, Jeddah, Saudi Arabia.
Background And Objective:
This study evaluates the capability of various general-purpose and healthcare-specialized Artificial Intelligence (AI) platforms in identifying Autism Spectrum Disorder (ASD) from clinical narratives.
Methods:
Using twenty standardized pediatric case reports (10 ASD and 10 non-ASD), the evaluation assessed diagnostic accuracy, concordance with clinical diagnoses, and statistical performance across different AI architectures.
Results:
The platforms demonstrated diverse operational profiles; Gemini 3 Pro achieved the highest rates of sensitivity and specificity, while other evaluated models exhibited sensitivity rates ranging from 60% to 90%. While statistical differences in performance between general-purpose and specialized systems were not significant (P> 0.799), advanced large language models showed the ability to reason through complex diagnostic narratives.
Conclusion:
These findings suggest that advanced general-purpose AI platforms can offer valuable support in interpreting complex ASD case narratives. To ensure clinical safety, incorporating stratified frameworks and standardized protocols remains essential. Further evaluation is required to determine the precise role of these tools.