Related Experiment Video
Updated: Mar 3, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
1.3K
Evaluating the Applicability of Advanced Large Language Models in Laboratory Medicine Test Questions: A Comparative
Wenzheng Han1, Wenkai Zhu1, Gang Feng1
1Department of Clinical Laboratory, The First Affiliated Hospital, Wannan Medical College, Wuhu, Anhui, People's Republic of China.
Advances in Medical Education and Practice
|March 2, 2026
Summary
Large language models (LLMs) show promise in medical laboratory science education. DeepSeek-R1 demonstrated strong performance, matching or exceeding human experts in accuracy and reasoning for complex questions.
Area of Science:
- Medical Laboratory Science
- Artificial Intelligence in Education
- Natural Language Processing
Background:
- Large language models (LLMs) show potential in medical education.
- Comprehensive performance assessment of LLMs in specialized fields like medical laboratory science is lacking.
Purpose of the Study:
- To evaluate advanced LLMs on medical laboratory science questions.
- Assess LLM accuracy, natural language generation (NLG) quality, reasoning, and efficiency.
Main Methods:
- Multi-faceted evaluation of DeepSeek-R1, Gemini-2.5 Pro, and GPT-5 against medical laboratory scientists.
- Utilized 493 knowledge- and reasoning-based questions (SCQs and MCQs) from a medical laboratory test bank.
- Measured performance via accuracy, Macro-F1, response time, NLG scores (ROUGE-L, METEOR), and logical reasoning assessment.
Main Results:
- DeepSeek-R1 achieved 78.3% accuracy, nearing senior expert performance (79.3%).
- DeepSeek-R1 excelled in complex reasoning-based MCQs (64.4% accuracy), outperforming human experts.
- DeepSeek-R1 demonstrated superior NLG quality and logical reasoning comprehensiveness, integrating crucial diagnostic findings.
Conclusions:
- DeepSeek-R1 shows potential to match or exceed senior expert performance in specific medical laboratory tasks.
- LLMs like DeepSeek-R1 can be effective tools for medical laboratory science education and assessment.
- Further research into LLM efficiency and integration into specialized scientific domains is warranted.
