Physicians and artificial intelligence diverge in evaluating large language models on real clinical cases

Peilun Shi1, Jian Li2, Ziqi Yang3

  • 1Department of Biomedical Engineering, The Chinese University of Hong Kong, Hong Kong SAR, China.

Summary

Evaluating large language models (LLMs) in healthcare requires real clinical cases. Physician assessments showed variability, indicating AI agents can assist but not replace human clinical judgment for medical applications.

Related Concept Videos