Large language models approach clinician performance in ESC cardiovascular risk stratification: a vignette-based

José Ferreira Santos1,2, Regina de Brito Duarte3,4, Inês Mota5

  • 1Católica Medical School, Sintra Campus, Estrada OctávioPato, 2635-631 Rio de Mouro, Lisboa, Portugal.

Insights

Large language models (LLMs) can extract cardiovascular risk factors from clinical text, but struggle with accurate risk score computation. The best LLMs matched clinician performance, highlighting computation as a key limitation.

Area of Science:

  • Cardiology
  • Artificial Intelligence
  • Clinical Decision Support

Background:

  • Cardiovascular risk stratification is crucial for patient management.
  • Guideline-based risk assessment requires data extraction, score computation, and category assignment.
  • Evaluating the capabilities of contemporary large language models (LLMs) in these tasks is essential.

Purpose of the Study:

  • To assess the performance of LLMs in cardiovascular risk stratification using the ESC SCORE2 framework.
  • To compare LLM performance against a benchmark of individual clinicians.
  • To identify specific areas of LLM limitations in the risk stratification process.

Main Methods:

  • Eleven LLMs were evaluated on 30 simulated clinical vignettes in Portuguese and English.
  • Models performed risk factor extraction, SCORE2 applicability determination, risk estimation, and risk category assignment.
  • Performance was compared against a reference standard set by expert cardiologists and an individual clinician benchmark.

Main Results:

  • LLMs demonstrated near-perfect risk factor extraction (micro-F1 0.97-0.99).
  • Agreement with expert-assigned risk categories was moderate (best LLM: κw 0.69), with a tendency to underestimate risk.
  • Post-hoc recalculation revealed computational execution as the primary failure mode, not data extraction.

Conclusions:

  • LLMs reliably extract cardiovascular risk information from clinical text.
  • The best-performing LLMs matched or exceeded average individual clinician performance in this structured task.
  • LLMs' main limitations are in downstream computation and rule application for risk stratification.
Abstract

Related Concept Videos