Related Experiment Video
Updated: Jan 10, 2026

Large Animal Model for Evaluating the Efficacy of the Gene Therapy in Ischemic Heart
Published on: September 2, 2021
Performance of large language models in interventional cardiology: the ILLUMINATE blinded model-comparison study
Attilio Lauretti1, Iginio Colaiori2, Simone Calcagno3
1Division of Cardiology, Santa Maria Goretti Hospital, Latina, Italy; Cardiology Unit, Department of Emergency and Admission, San Paolo Hospital, Civitavecchia, Italy; Department of Cardiovascular Sciences, Fondazione Policlinico Agostino Gemelli IRCCS, Rome, Italy; Division of Cardiology, Cardiovascular and Thoracic Department, Città della Salute e della Scienza, Turin, Italy; Division of Cardiology, Department of Medical Sciences, University of Turin, Italy; Department of Medical-Surgical Sciences and Biotechnologies, Sapienza University of Rome, Latina, Italy; Maria Cecilia Hospital, GVM Care and Research, Cotignola, Italy; ICOT Marco Pasquali Institute, Cardiovascular Department Latina; Department of Clinical and Molecular Medicine, Sapienza University of Rome, Rome, Italy.
Large language models (LLMs) show potential for interventional cardiology (IC) decision-making, but performance varies. ChatGPT with internet search or guidelines excelled, while Gemini performed lowest, indicating a need for improved LLM integration.
Area of Science:
- Cardiology
- Artificial Intelligence
- Medical Informatics
Background:
- Large language models (LLMs) offer potential for complex decision-making in interventional cardiology (IC).
- However, their comparative effectiveness in providing clinical recommendations is not well-established.
- This study evaluates and compares the quality of recommendations from six LLMs for complex IC cases.
Purpose of the Study:
- To compare the performance of six different large language models (LLMs) in generating clinical recommendations for interventional cardiology (IC).
- To assess the quality of LLM-generated recommendations based on appropriateness, accuracy, relevance, clarity, and clinical utility.
Main Methods:
- Twenty complex interventional cardiology cases (10 coronary artery disease, 10 structural heart disease) were developed.
- Six LLMs were evaluated: ChatGPT (default, guideline-integrated, internet-enabled), Gemini, Mistral 7B, and Perplexity AI.
- Five interventional cardiology experts independently assessed anonymized LLM outputs using a 0-10 scale, with statistical analysis via a mixed linear model.
Main Results:
- Overall composite score was 7.1, with significant performance variation across LLMs (P < .001).
- ChatGPT with internet search (7.8) and guideline integration (7.7) significantly outperformed other models.
- Gemini scored lowest (6.3), while ChatGPT default, Mistral 7B, and Perplexity AI showed moderate performance (6.9-7.0).
Conclusions:
- LLMs demonstrate promise for interventional cardiology decision support but are currently suboptimal.
- Integrating web search and guideline-based knowledge retrieval is crucial for maximizing LLM utility.
- Further development is needed to enhance the reliability and clinical applicability of LLMs in cardiology.

