Related Experiment Video
Updated: May 23, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
DeepSeek vs ChatGPT vs Claude: benchmarking large language models for clinical diagnosis using a novel
Lichen Du1, Jiachen Zhong2, Shirong Zheng3
1Department of Biostatistics, University of Southern California, Los Angeles, CA, 90033, USA.
BMC Medical Informatics and Decision Making
|May 22, 2026
Summary
State-of-the-art large language models (LLMs) show promise in clinical diagnosis, with DeepSeek-V3 performing best in an open-ended diagnostic task benchmark. The study utilized a novel ICD-10-CM-based framework for reproducible evaluation.
Area of Science:
- Artificial Intelligence in Medicine
- Clinical Decision Support Systems
- Natural Language Processing in Healthcare
Background:
- Evaluating large language models (LLMs) for clinical reasoning is crucial but challenging due to limitations in existing benchmarks.
- Traditional methods often use multiple-choice questions or subjective ratings, failing to capture real-world diagnostic complexity.
- A need exists for clinically grounded, reproducible criteria to assess LLM performance in open-ended diagnostic tasks.
Purpose of the Study:
- To benchmark state-of-the-art LLMs for open-ended clinical diagnosis.
- To introduce and utilize a hierarchical International Classification of Diseases, Tenth Revision, Clinical Modification (ICD-10-CM)-based evaluation framework.
- To assess LLM diagnostic accuracy and consistency on real-world clinical cases.
Main Methods:
- Four LLMs (DeepSeek-V3, GPT-4o, GPT-4o mini, Claude 3.5 Sonnet) were evaluated on 50 recent clinical cases from the MultiCaRe dataset.
- Models generated ranked differential diagnoses and ICD-10-CM codes for each case.
- A hierarchical scoring framework based on ICD-10-CM structure assessed diagnostic performance and specificity, with two runs per case for consistency.
Main Results:
- DeepSeek-V3 achieved the highest mean diagnostic score (2.32 ± 1.53), closely followed by GPT-4o (2.23 ± 1.50).
- Binary adjusted diagnostic accuracy ranged from 76% (Claude 3.5 Sonnet) to 82% (DeepSeek-V3, GPT-4o).
- Models performed better at identifying disease categories than specific diagnoses; DeepSeek-V3 demonstrated superior stability.
Conclusions:
- State-of-the-art LLMs exhibit significant potential for open-ended clinical diagnosis, particularly in identifying relevant disease categories.
- DeepSeek-V3 demonstrated the strongest numerical performance in this benchmark, though overall model differences were modest.
- The proposed ICD-10-CM-based hierarchical framework offers a reproducible and clinically relevant method for evaluating LLM diagnostic capabilities.