Related Experiment Video
Updated: May 26, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Utility of large language models as information tools for nursing care in gout: a comparative study of DeepSeek and
Xia Pan1, Yali Wang1, QiaoLan Yang2
1Department of Rheumatology and Immunology, The Affiliated Guangdong Second Provincial GeneralHospital of Jinan University, Guangzhou, China.
Background:
With the rapid advancement of artificial intelligence, LLMs (LLMs) are now employed across diverse domains. In nursing, their capacity for high-quality content generation is especially promising, offering practical value for clinical management, research, and education. Among the leading Chinese models is DeepSeek-R1.
Objective:
This study aims to evaluate and compare the effectiveness of DeepSeek-R1 and ChatGPT-4.0 as online information sources for nursing professionals seeking evidence-based care strategies for gout patients.
Methods:
We identified the 15 highest-priority questions on gout and related nursing strategies by surveying the research site, patients, and healthcare providers. These questions, posed in Chinese, were separately submitted to DeepSeek-R1 and ChatGPT-4.0. The Flesch Kincaid Grade Level (FKGL) and the Flesch Reading Ease (FRE) were used to evaluate the readability of their answers. The mDISCERN score was employed to compare the accuracy of their responses, and the age of statistical reference materials was assessed to compare their timeliness. GraphPad Prism 8.0.1 was used for all statistical analyses and figure preparation.
Results:
Readability and citation characteristics differed between the two LLMs. The FKGL of DeepSeek-R1 (13.04 ± 1.62) exceeded that of ChatGPT-4.0 (11.41 ± 1.74; p = 0.013), whereas FRE was lower for DeepSeek-R1 (40.50 ± 8.12) than for ChatGPT-4.0 (49.08 ± 8.90; p = 0.010). The mDISCERN quality score was numerically higher for ChatGPT-4.0 (4.30 ± 0.73) than for DeepSeek-R1 (3.98 ± 0.70), but this difference was not statistically significant (p = 0.16). DeepSeek cited 21 sources and ChatGPT-4.0 23; clinical guidelines predominated in both corpora (38.1 vs. 47.8 %, respectively). The mean publication age (years elapsed from 2025) was significantly younger for DeepSeek-R1 (3.57 ± 2.33) than for ChatGPT-4.0 (5.42 ± 2.34; p < 0.05). In addition, DeepSeek-R1 provided 4 reference links were invalid.
Conclusion:
Both DeepSeek-R1 and ChatGPT-4.0 drew chiefly from high-level evidence and produced accurate, professional answers; ChatGPT-4.0 rendered them in markedly clearer prose. While DeepSeek-R1 offered more up-to-date citations, several of its reference links were non-functional.
