Related Experiment Video
Updated: Jun 29, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Chain-of-verification prompting for NIH stroke scale extraction using small and frontier large language models
Jack Lott1, Brett Stubbert2, David McShannon3
1School of Medicine, Queen's University, 80 Barrie Street, Kingston, ON K7L 3N6, Canada.
Background:
The National Institutes of Health Stroke Scale (NIHSS) is critical to acute stroke care but is often documented in unstructured notes. Large language models (LLMs) can enable automated extraction, though smaller models often underperform relative to frontier systems. Chain-of-Verification (CoVe) prompting introduces a structured self-verification step that may improve performance.
Methods:
We evaluated eight LLMs on 312 discharge summaries. Small models included LLaMA 3.2 3B, Ministral 3B, Gemma 3 4B, and Qwen 3 4B. Frontier models included GPT-5.2, Gemini 3 Pro, Claude Opus 4.5, and Grok 4. Each model was tested under a baseline and CoVe prompt. Outcomes were subscore exact-match accuracy, subscore mean absolute error (MAE), total score exact-match accuracy, and total score MAE.
Results:
At baseline, small models achieved 53.2 ± 10.0% subscore accuracy and subscore MAE 0.84 ± 0.22, compared with 88.5 ± 10.1% and 0.15 ± 0.16 in frontier models (both p < 0.001). Total exact accuracy was low in both groups (7.7 ± 12.9% vs 35.9 ± 32.4%). CoVe significantly improved small-model performance (subscore accuracy 65.0 ± 10.9%; subscore MAE 0.55 ± 0.21; total MAE 4.84 ± 2.30 vs 7.19 ± 3.54 at baseline; all p < 0.001), although total exact accuracy remained modest (9.6 ± 15.7%). Frontier models showed no significant group-level change with CoVe.
Conclusion:
CoVe prompting substantially improves NIHSS extraction in small LLMs while producing negligible effects in frontier models. Although smaller model performance remains insufficient for standalone clinical deployment, CoVe prompting offers a promising avenue for further exploration.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy