Related Experiment Video
Updated: Jun 16, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Blinded two-phase evaluation of large language models in complex cardiac surgery: task-specific performance and
Marc Leon1, Ruibin Feng2, Manuel Quiroz Flores1
1Department of Cardiothoracic Surgery, Stanford University School of Medicine, Stanford, CA, United States.
Frontiers in Digital Health
|June 15, 2026
Summary
Large language models (LLMs) show promise but are not yet safe for complex surgical decisions. Human-LLM collaboration reveals overacceptance of flawed AI reasoning, highlighting current limitations.
Area of Science:
- Artificial Intelligence in Medicine
- Surgical Decision Support Systems
- Human-Computer Interaction in Healthcare
Background:
- Large language models (LLMs) demonstrate proficiency in medical benchmarks but their utility in complex surgical decision-making remains unevaluated.
- The integration of LLM-generated reasoning into clinical practice, particularly concerning clinician recognition and acceptance, is an under-explored area.
- Assessing LLM performance and human-LLM collaboration is crucial for safe implementation in critical care settings.
Purpose of the Study:
- To develop and implement a novel two-phase evaluation framework for assessing LLM performance in cardiac surgery.
- To investigate the dynamics of human-LLM collaboration, focusing on clinicians' ability to critically evaluate AI-generated reasoning.
- To identify specific limitations of current LLMs in complex surgical scenarios and inform future development.
Main Methods:
- Creation of 15 high-fidelity cardiac surgery scenarios with expert-defined reasoning tasks and reference answers.
- Evaluation of five leading LLMs (O1, O3-mini-high, DeepSeek-R1, GPT-4, Llama3-OpenBioLLM-70B) using a multi-agent prompting strategy.
- A blinded, two-phase evaluation by senior cardiac surgeons assessing LLM performance independently and after reviewing reference answers to gauge judgment shifts.
Main Results:
- LLM performance varied, with O1 achieving the highest median normalized score (0.896).
- Models struggled with patient safety (0.507), hallucination avoidance (0.549), and clinical efficiency (0.597) dimensions.
- Clinician overacceptance of AI reasoning was observed, with a significant percentage of second-round ratings revised negatively, indicating collaboration imbalances.
Conclusions:
- Reasoning-optimized LLMs showed superior performance but all models exhibited clinical limitations, particularly in complex, longitudinal reasoning.
- Current LLMs are not sufficiently reliable for safe integration into complex surgical decision-making due to performance gaps.
- Addressing the human-LLM collaboration imbalance, specifically mitigating overacceptance of incorrect AI reasoning, is critical for future clinical deployment.