Blinded two-phase evaluation of large language models in complex cardiac surgery: task-specific performance and

Marc Leon1, Ruibin Feng2, Manuel Quiroz Flores1

  • 1Department of Cardiothoracic Surgery, Stanford University School of Medicine, Stanford, CA, United States.

Summary

Large language models (LLMs) show promise but are not yet safe for complex surgical decisions. Human-LLM collaboration reveals overacceptance of flawed AI reasoning, highlighting current limitations.

Related Concept Videos