Related Experiment Video
Updated: Aug 14, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
How Often Do Large Language Models Agree with Each Other-And with the Truth? A Consensus- and Complexity-Stratified
Nafiye Sanlier1, Umid Sulaimanov1, Ariorad Moniri2
1Department of Neurological Surgery, School of Medicine and Public Health, University of Wisconsin-Madison, Madison, WI 53792, USA.
Journal of Clinical Medicine
|August 13, 2026
Summary
Inter-model consensus aids large language model (LLM) data extraction in neuroimaging AI, but complexity matters. A hybrid approach can significantly reduce review effort by automating simple variables and verifying complex ones.
Area of Science:
- Neuroimaging AI
- Natural Language Processing
- Biomedical Informatics
Background:
- Reliable integration of large language models (LLMs) into neuroimaging data extraction workflows is challenging.
- Existing benchmarks may underestimate LLM performance by relying solely on exact-match accuracy.
- The utility of inter-model consensus and variable complexity for guiding automated extraction remains unclear.
Purpose of the Study:
- To evaluate inter-model consensus as a confidence signal for human-AI neuroimaging data extraction.
- To assess if variable complexity can guide automated workflow triage in neuroimaging AI.
- To compare LLM extraction performance against expert references using both exact-match and semantic-equivalence accuracy.
Main Methods:
- Four advanced LLMs were prompted to extract 22 variables from 91 neuroimaging AI articles.
- Variables were categorized into low, medium, and high complexity.
- Item-level consensus and five triage strategies were analyzed for efficiency-accuracy trade-offs.
Main Results:
- Semantic-equivalence accuracy ranged from 80.5-83.4% across models.
- Unanimous consensus (4/4 models) achieved 85.8% exact-match accuracy, improving to 95.3% after semantic normalization.
- Extraction reliability was strongly dependent on variable complexity (96.6% for low, 73.2% for medium, 38.1% for high).
- A hybrid strategy reduced review effort by approximately 59%.
Conclusions:
- Inter-model consensus is a valuable but incomplete signal for LLM-assisted neuroimaging data extraction.
- Workflow automation must be complexity-stratified: automate low-complexity, verify medium-complexity, and maintain human oversight for high-complexity variables.
- LLM integration into neuroimaging AI extraction requires tailored, complexity-aware workflow design.
