Related Experiment Video
Updated: Jun 13, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating Large Language Models for Automated Evidence Synthesis in Neuroimaging AI: A Multi-Model Benchmark
Umid Sulaimanov1, Nafiye Sanlier1, Ariorad Moniri2
1Department of Neurological Surgery, School of Medicine and Public Health, University of Wisconsin-Madison, Madison, WI 53792, USA.
Journal of Clinical Medicine
|June 12, 2026
Summary
Large language models (LLMs) show promise for automating data extraction in systematic reviews, but struggle with complex neuroimaging AI literature. Gemini 3 Pro Preview led in accuracy, though human oversight remains crucial for nuanced data.
Area of Science:
- Artificial Intelligence
- Neuroimaging
- Systematic Reviews
Background:
- Data extraction for systematic reviews is time-consuming and resource-intensive.
- Evaluating the utility of advanced AI in automating evidence synthesis is critical.
- Specialized neuroimaging artificial intelligence (AI) literature presents unique challenges for data extraction.
Purpose of the Study:
- To assess the performance of four leading large language models (LLMs) in extracting structured metadata from neuroimaging AI literature.
- To compare the accuracy of Google Gemini 3 Pro Preview, Anthropic Claude Opus 4.5, Perplexity Sonar Pro, and OpenAI GPT 5.2 for complex data extraction tasks.
- To determine the impact of variable complexity on LLM performance in automated evidence synthesis.
Main Methods:
- A standardized prompt was used to extract 22 variables from 91 neuroimaging AI articles.
- Variables were categorized into low, medium, and high complexity tiers.
- Performance was evaluated using exact-match accuracy against expert-validated ground truth.
Main Results:
- Gemini 3 Pro Preview achieved the highest overall exact-match accuracy (56.4%), outperforming other models.
- Model performance decreased significantly with increasing variable complexity.
- Accuracy for low-complexity fields was high (88.9-92.9%), while high-complexity variables yielded very low accuracy (2.7-15.5%).
Conclusions:
- Frontier LLMs can automate the extraction of simple, categorical data effectively.
- Complex methodological variables requiring clinical judgment or multi-section synthesis remain challenging for current LLMs.
- Human review is indispensable for ensuring accuracy in extracting context-dependent variables from specialized literature.
