Assessing DeepSeek-R1 for Clinical Decision Support in Multidisciplinary Laboratory Medicine
Qinpeng Li1, Lili Zhan1, Xinjian Cai1
1Department of Clinical Laboratory Medicine, National Cancer Center/National Clinical Research Center for Cancer/Cancer Hospital & Shenzhen Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College, Shenzhen, Guangdong, People's Republic of China.
Journal of Multidisciplinary Healthcare
|August 18, 2025
Summary
Artificial intelligence (AI) using large language models (LLMs) shows promise in clinical laboratory medicine. DeepSeek-R1 achieved 72.9% accuracy in diagnostic hypotheses but needs improvement for differential diagnoses.
Area of Science:
- Artificial Intelligence in Medicine
- Clinical Laboratory Diagnostics
- Medical Decision Support
Background:
- Advancements in artificial intelligence (AI), specifically large language models (LLMs), are revolutionizing healthcare.
- LLMs like DeepSeek-R1 offer potential for enhanced diagnostic decision-making and clinical workflow optimization in laboratory medicine.
- The integration of AI aims to improve diagnostic accuracy, support clinical judgment, and refine patient care strategies.
Purpose of the Study:
- To evaluate the performance of the DeepSeek-R1 large language model in analyzing clinical laboratory cases.
- To assess the accuracy and completeness of DeepSeek-R1 in generating diagnostic hypotheses, differential diagnoses, and diagnostic workups.
- To determine the utility of DeepSeek-R1 as a decision-support tool across diverse clinical scenarios in laboratory medicine.
Main Methods:
- Analysis of 100 clinical laboratory cases from "Clinical Laboratory Medicine Case Studies."
- Independent querying of DeepSeek-R1 three times per case for diagnosis, differential diagnoses, and diagnostic tests.
- Assessment of AI-generated outputs for accuracy and completeness by senior clinical laboratory physicians.
Main Results:
- DeepSeek-R1 achieved an overall accuracy of 72.9% and completeness of 73.4%.
- Highest accuracy was observed for diagnostic hypotheses (85.7%), while differential diagnoses showed lower accuracy (55.0%).
- Performance varied by disease category, with superior results in genetic and obstetric diagnostics (accuracy 93.1%).
Conclusions:
- DeepSeek-R1 shows potential as a decision-support tool for diagnostic hypotheses and workup recommendations in clinical laboratory medicine.
- Limitations exist in differential diagnosis generation and handling clinical nuances, requiring further development.
- Future research should focus on expanding training data and incorporating physician feedback for enhanced real-world applicability.


