Related Experiment Video
Updated: May 2, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
1.3K
Open-Source Large Language Models Distilled DeepSeek-R1 Pose Challenges for On-Premises Clinical Deployment in
Wei Zhong1, Yiyao Fu2, Dingchuan Peng3
1Department of Prenatal Diagnosis, Beijing Obstetrics and Gynecology Hospital, Capital Medical University. Beijing Maternal and Child Health Care Hospital, Beijing, 100026, China.
Journal of Medical Systems
|April 30, 2026
Summary
The large DeepSeek-R1-671B model shows promise for medical diagnosis, but distilled versions do not outperform their base models. Further validation is needed before deploying distilled models for clinical use.
Area of Science:
- Artificial Intelligence
- Medical Informatics
- Computational Linguistics
Background:
- Open-source reasoning large language models (LLMs) like DeepSeek-R1 are entering clinical settings.
- The diagnostic performance of various DeepSeek-R1 parameter versions, particularly distilled models, requires thorough evaluation.
Purpose of the Study:
- To compare the diagnostic accuracy of five DeepSeek-R1 models against their respective base models.
- To assess the performance of distilled LLMs in simulated clinical diagnostic tasks.
Main Methods:
- Paired comparisons of five DeepSeek-R1 models (including distilled versions) with their base models.
- Testing on 110 simulated clinical cases across internal medicine, surgery, neurology, gynecology, and pediatrics.
- Accuracy assessment using McNemar's test with a significance threshold of 0.01.
Main Results:
- DeepSeek-R1-671B significantly outperformed its base model (DeepSeek-V3).
- DeepSeek-R1-8B (distilled) underperformed compared to its base model (Llama3.1-8B).
- No significant performance differences were found for mid-sized models; common error modes in distilled models included reasoning drift and diagnostic priority inversion.
Conclusions:
- The largest DeepSeek-R1 model (671B) demonstrates potential for medical diagnosis.
- Distilled DeepSeek-R1 models do not show superior diagnostic performance compared to their larger counterparts.
- Current findings do not support the deployment of distilled models for text-based diagnosis without real-world data validation.
