Related Experiment Video
Updated: May 31, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large language models approach clinician performance in ESC cardiovascular risk stratification: a vignette-based
José Ferreira Santos1,2, Regina de Brito Duarte3,4, Inês Mota5
1Católica Medical School, Sintra Campus, Estrada OctávioPato, 2635-631 Rio de Mouro, Lisboa, Portugal.
Insights
Large language models (LLMs) can extract cardiovascular risk factors from clinical text, but struggle with accurate risk score computation. The best LLMs matched clinician performance, highlighting computation as a key limitation.
Area of Science:
- Cardiology
- Artificial Intelligence
- Clinical Decision Support
Background:
- Cardiovascular risk stratification is crucial for patient management.
- Guideline-based risk assessment requires data extraction, score computation, and category assignment.
- Evaluating the capabilities of contemporary large language models (LLMs) in these tasks is essential.
Purpose of the Study:
- To assess the performance of LLMs in cardiovascular risk stratification using the ESC SCORE2 framework.
- To compare LLM performance against a benchmark of individual clinicians.
- To identify specific areas of LLM limitations in the risk stratification process.
Main Methods:
- Eleven LLMs were evaluated on 30 simulated clinical vignettes in Portuguese and English.
- Models performed risk factor extraction, SCORE2 applicability determination, risk estimation, and risk category assignment.
- Performance was compared against a reference standard set by expert cardiologists and an individual clinician benchmark.
Main Results:
- LLMs demonstrated near-perfect risk factor extraction (micro-F1 0.97-0.99).
- Agreement with expert-assigned risk categories was moderate (best LLM: κw 0.69), with a tendency to underestimate risk.
- Post-hoc recalculation revealed computational execution as the primary failure mode, not data extraction.
Conclusions:
- LLMs reliably extract cardiovascular risk information from clinical text.
- The best-performing LLMs matched or exceeded average individual clinician performance in this structured task.
- LLMs' main limitations are in downstream computation and rule application for risk stratification.
Aims:
Guideline-based cardiovascular risk stratification requires three distinct competencies: extracting risk factor data from clinical text, computing a validated risk score, and applying guideline-defined thresholds to assign a final risk category. We evaluated contemporary large language models (LLMs) on each of these tasks within the European Society of Cardiology (ESC) SCORE2 framework and compared LLM performance against a pooled individual clinician benchmark to contextualize findings against real-world human reproducibility.
Methods And Results:
Eleven LLMs were evaluated using 30 simulated outpatient clinical vignettes presented in both Portuguese and English. For each vignette, models extracted cardiovascular risk factors, determined SCORE2 applicability, generated 10-year risk estimates where appropriate, and assigned a final three-class ESC risk category. A committee of three cardiologists established the reference standard; eight independent clinicians provided an individual-level human benchmark. Traditional risk-factor extraction was near-perfect across all models (micro-F1 0.97-0.99). Agreement with expert-assigned final risk categories was moderate and variable (best: GPT-4o, quadratic-weighted κw 0.69, 95% CI 0.44-0.84), with 10 of 11 models more often underestimating than overestimating risk. To isolate the source of classification error, post hoc deterministic recalculation of SCORE2 was performed using model-extracted variables in eligible vignettes; this markedly improved agreement across all models (κw 0.85-0.90), demonstrating that extraction was largely intact and computational execution was the primary failure mode. The pooled individual clinician benchmark showed moderate agreement with the reference standard (κw 0.52, 95% CI 0.28-0.67), indicating that the best-performing LLMs matched or exceeded the average individual clinician on this guideline-based task. Performance was broadly consistent across Portuguese and English.
Conclusion:
Contemporary LLMs reliably extract cardiovascular risk information from clinical text, and the best-performing systems achieved agreement within the range of average individual clinicians on this structured task. Their principal limitation lies in downstream computation and rule application.