DeepSeek vs ChatGPT vs Claude: benchmarking large language models for clinical diagnosis using a novel

Lichen Du1, Jiachen Zhong2, Shirong Zheng3

  • 1Department of Biostatistics, University of Southern California, Los Angeles, CA, 90033, USA.

Summary

State-of-the-art large language models (LLMs) show promise in clinical diagnosis, with DeepSeek-V3 performing best in an open-ended diagnostic task benchmark. The study utilized a novel ICD-10-CM-based framework for reproducible evaluation.