Multimodal LLM vs. Human-Measured Features for AI Predictions of Autism in Home Videos
Parnian Azizian1, Mohammadmahdi Honarmand1, Aditi Jaiswal2
1Department of Mechanical Engineering, Stanford University, Stanford, CA 94305, USA.
Summary
Large language models (LLMs) show promise for autism detection, achieving up to 90% accuracy in recent evaluations. This AI advancement could improve early diagnosis accessibility and scalability for children.
Area of Science:
- Artificial Intelligence in Healthcare
- Neurodevelopmental Disorder Diagnostics
- Machine Learning for Medical Applications
Background:
- Autism diagnosis faces challenges, with average ages of 5-8, particularly in underserved populations.
- Current AI diagnostics (e.g., CanvasDx) utilize a human-in-the-loop approach, but fully automated methods for early detection remain underdeveloped.
- Early autism detection before age 3 is crucial, necessitating scalable, accessible, and privacy-preserving AI solutions.
Purpose of the Study:
- To systematically evaluate multimodal Large Language Models (LLMs) as direct replacements for human annotators in AI-based autism detection.
- To assess the performance progression of different Gemini model variants in classifying autism from video data.
- To compare LLM performance against clinical baselines and human annotator capabilities.
Main Methods:
- Seven Gemini model variants (1.5-2.5 series) were evaluated on 50 YouTube videos.
- Performance was measured using classification accuracy against validated autism detection AI models (LR5).
- LLM assessment strategies were compared with those of crowdworkers and clinicians.
Main Results:
- Gemini models demonstrated significant generational improvement, with version 2.5 achieving 85-90% accuracy (best: 89.6% for 2.5 Pro).
- LLM performance approached clinical baselines (88%) and crowdworker performance (92-98%).
- LLMs showed high internal consistency, focusing on language/behavioral markers, contrasting with human annotators' strategies.
Conclusions:
- Multimodal LLMs show rapid advancement and potential as automated tools for autism detection.
- While not yet matching top human annotators in feature extraction for human-in-the-loop systems, LLMs offer advantages in consistency, scalability, cost, and privacy.
- LLMs are positioned as potentially viable future alternatives to aid in the autism diagnostic process.
