Related Experiment Video
Updated: Jun 6, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
A multi-agent large language model framework to automatically assess performance of a clinical AI Triage tool
Adam E Flanders1, Yifan Peng2, Luciano Prevedello3
1Thomas Jefferson University, Philadelphia, PA, USA.
Npj Health Systems
|June 5, 2026
Summary
An ensemble of open-source large language models (LLMs) reliably evaluated clinical AI triage tools for intracranial hemorrhage (ICH) detection. Combining multiple LLMs offers a consistent method for retrospective AI performance assessment.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Informatics
- Natural Language Processing
Background:
- Radiology reports are valuable for assessing clinical artificial intelligence (AI) tool performance.
- Evaluating AI tools retrospectively requires reliable ground truth data.
Purpose of the Study:
- To assess the efficacy of an ensemble of open-source large language models (LLMs) in evaluating clinical AI triage tools for detecting intracranial hemorrhage (ICH).
- To compare the performance of LLM ensembles against individual LLMs and human review for retrospective AI evaluation.
Main Methods:
- Eight open-source LLMs and a proprietary LLM (GPT-4o) analyzed radiology reports for ICH presence using a multi-shot prompt.
- Performance was evaluated by comparing LLM outputs to human report review.
- Consensus from LLM ensembles (Full-9, Top-3) was compared to individual LLM performance and human assessment.
Main Results:
- Individual LLM capabilities varied, with llama3.3:70b and GPT-4o achieving the highest Area Under the Curve (AUC).
- The Full-9 Ensemble, Top-3 Ensemble, and consensus methods demonstrated strong performance using Matthews Correlation Coefficient (MCC).
- No statistically significant performance differences were found between the Top-3 Ensemble, Full-9 Ensemble, and consensus methods.
Conclusions:
- An ensemble of open-source LLMs provides a consistent and reliable method for retrospective ground truth evaluation of clinical AI triage tools.
- Ensemble approaches mitigate the variability of individual LLM performance, enhancing evaluation reliability.
- LLM ensembles offer a viable alternative to single LLM or manual review for retrospective AI performance assessment.
