Evaluating Large-Language Models Against Providers on Surgical Diagnostic Reasoning Tasks

Taj Keshav1, David Chow2, Tiffany Kippenberger2

  • 1Department of Preventive Medicine and Biostatistics, Uniformed Services University of the Health Sciences, Bethesda, Maryland.

Summary

General surgery residents demonstrated superior accuracy in differential diagnoses compared to large language models (LLMs). LLMs showed higher internal consensus, suggesting context-specific utility in surgical education.

Related Concept Videos