Related Experiment Video
Updated: May 12, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
462
Benchmark evaluation of DeepSeek large language models in clinical decision-making
Sarah Sandmann1, Stefan Hegselmann2, Michael Fujarski1
1Institute of Medical Informatics, University of Münster, Münster, Germany.
Nature Medicine
|April 23, 2025
Summary
Open-source large language models (LLMs) like DeepSeek match proprietary LLMs in clinical decision support. This offers a secure, compliant path for AI in healthcare.
Area of Science:
- Artificial Intelligence in Medicine
- Clinical Informatics
- Medical Decision Support Systems
Background:
- Proprietary large language models (LLMs) face adoption barriers in healthcare due to privacy regulations.
- Open-source LLMs offer potential for on-site deployment and data customization in hospitals.
Purpose of the Study:
- To benchmark the clinical utility of open-source DeepSeek models against proprietary LLMs.
- To evaluate performance on clinical decision support tasks using real-world patient cases.
Main Methods:
- Performance comparison of DeepSeek-V3 and DeepSeek-R1 against GPT-4o and Gemini-2.0 Flash Thinking Experimental.
- Benchmarking conducted on 125 patient cases with diverse disease profiles.
- Evaluation focused on clinical decision support capabilities.
Main Results:
- DeepSeek models demonstrated comparable or superior performance to proprietary LLMs.
- Open-source LLMs achieved high efficacy in clinical decision support tasks.
- No significant performance difference was observed between open-source and proprietary models.
Conclusions:
- Open-source LLMs are a viable and effective alternative for clinical decision support.
- These models facilitate secure, compliant, and scalable AI implementation in healthcare settings.
- DeepSeek models present a promising pathway for advancing medical AI applications.

