Related Experiment Video
Updated: Aug 13, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Critical Care-Specific vs General-Purpose Large Language Models in Emergency Intensive Care Unit Diagnosis:
Lihong Zheng1,2, Zeyu Lin1,2, Xiaolu Liu3
1Department of Emergency Medicine, Peking University Shenzhen Hospital, Shenzhen, Guangdong, China.
Journal of Medical Internet Research
|August 11, 2026
Summary
A specialized large language model (LLM) for critical care showed comparable diagnostic accuracy to general LLMs in the emergency intensive care unit (EICU). Further validation is needed before clinical deployment as an independent diagnostic tool.
Area of Science:
- Critical care medicine
- Artificial intelligence in healthcare
- Medical diagnostics
Background:
- Emergency intensive care units (EICUs) manage critically ill patients, where diagnostic errors are frequent and have severe consequences.
- Large language models (LLMs) show promise as decision-support tools, but comparative data for critical care is limited.
Purpose of the Study:
- To compare the diagnostic accuracy of a critical care-specific LLM (Qiyuan 3.0.1) against general-purpose LLMs (GPT-5.1, DeepSeek V.3.1, Qwen3-32B) for EICU diseases.
- To provide evidence for selecting appropriate AI tools in critical care settings.
Main Methods:
- A retrospective study of 184 EICU patients.
- Two datasets were created: initial (24 hours) and final (complete clinical course).
- Four LLMs were evaluated using zero-shot prompts, with diagnoses compared against a consensus of expert intensivists.
Main Results:
- The critical care LLM (Qiyuan 3.0.1) achieved 64.1% top-1 accuracy, comparable to top general LLMs (GPT-5.1 at 59.2%, DeepSeek V3.1 at 57.1%).
- All models performed significantly better than Qwen3-32B (51.6%), but no significant differences were found between Qiyuan 3.0.1 and the top two general models.
- Model performance varied across different EICU diseases, and all models had top-1 accuracy below 70%.
Conclusions:
- The critical care LLM demonstrated comparable performance to leading general LLMs in this EICU study.
- Current accuracy levels are insufficient for independent clinical deployment.
- Future steps include external validation, human-machine collaboration trials, and model optimization for critical care applications.