Related Experiment Video
Updated: May 31, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large language models approach clinician performance in ESC cardiovascular risk stratification: a vignette-based
José Ferreira Santos1,2, Regina de Brito Duarte3,4, Inês Mota5
1Católica Medical School, Sintra Campus, Estrada OctávioPato, 2635-631 Rio de Mouro, Lisboa, Portugal.
European Heart Journal. Digital Health
|May 29, 2026
Summary
Large language models (LLMs) can extract cardiovascular risk factors from clinical text, but struggle with accurate risk score computation. The best LLMs matched clinician performance, highlighting computation as a key limitation.
Area of Science:
- Cardiology
- Artificial Intelligence
- Clinical Decision Support
Background:
- Cardiovascular risk stratification is crucial for patient management.
- Guideline-based risk assessment requires data extraction, score computation, and category assignment.
- Evaluating the capabilities of contemporary large language models (LLMs) in these tasks is essential.
Purpose of the Study:
- To assess the performance of LLMs in cardiovascular risk stratification using the ESC SCORE2 framework.
- To compare LLM performance against a benchmark of individual clinicians.
- To identify specific areas of LLM limitations in the risk stratification process.
Main Methods:
- Eleven LLMs were evaluated on 30 simulated clinical vignettes in Portuguese and English.
- Models performed risk factor extraction, SCORE2 applicability determination, risk estimation, and risk category assignment.
- Performance was compared against a reference standard set by expert cardiologists and an individual clinician benchmark.
Main Results:
- LLMs demonstrated near-perfect risk factor extraction (micro-F1 0.97-0.99).
- Agreement with expert-assigned risk categories was moderate (best LLM: κw 0.69), with a tendency to underestimate risk.
- Post-hoc recalculation revealed computational execution as the primary failure mode, not data extraction.
Conclusions:
- LLMs reliably extract cardiovascular risk information from clinical text.
- The best-performing LLMs matched or exceeded average individual clinician performance in this structured task.
- LLMs' main limitations are in downstream computation and rule application for risk stratification.