Evaluating Large Language Models for Automated Evidence Synthesis in Neuroimaging AI: A Multi-Model Benchmark

Umid Sulaimanov1, Nafiye Sanlier1, Ariorad Moniri2

  • 1Department of Neurological Surgery, School of Medicine and Public Health, University of Wisconsin-Madison, Madison, WI 53792, USA.

Summary

Large language models (LLMs) show promise for automating data extraction in systematic reviews, but struggle with complex neuroimaging AI literature. Gemini 3 Pro Preview led in accuracy, though human oversight remains crucial for nuanced data.

Related Concept Videos