Related Experiment Video
Updated: Apr 7, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Accuracy and readability of LLM-generated responses for gout management: a benchmark study based on ACR guidelines
Zhuhai Shao1, Yiwen Zhang2, Bingfei Cheng1
1Department of Endocrine and Metabolic Diseases, the Affiliated Hospital of Qingdao University, Qingdao, China.
Background:
Gout and hyperuricemia, linked to purine metabolism abnormalities or impaired uric acid excretion, are rising with lifestyle changes. Effective self-management and health literacy are crucial for gout management. While large language models (LLMs) show promise in enhancing health management, their potential on gout patient education remains underexplored. This study aimed to evaluate the accuracy and readability of responses generated by three LLMs, including DeepSeek-V3, DeepSeek-R1, and GPT-5, to questions based on the gout and hyperuricemia guidelines published by the American College of Rheumatology (ACR).
Methods:
Based on the ACR gout guidelines, a set of 42 questions was curated and submitted to three LLMs. Their responses were independently rated on a 5-point Likert scale by three expert gout and hyperuricemia specialists against the guidelines. Accuracy was rated as either an average score of ≥ 4 (low threshold) or 5 (high threshold). The readability of the responses was assessed by Microsoft Word, which provided metrics including the word count, character count, Flesch Reading Ease (FRE) score, Flesch-Kincaid Grade Level (FKGL), and Automated Readability Index (ARI) were calculated.
Results:
Our findings reveal that response accuracy was significantly higher for GPT-5 compared to DeepSeek-V3 (P < 0.001), with no significant difference between GPT-5 and DeepSeek-R1 (P > 0.05). In terms of readability, GPT-5 produced the most complex responses (FKGL: 12.89 ± 2.22, ARI: 14.87 ± 2.40), while DeepSeek-R1 generated the longest outputs.
Conclusion:
LLMs show potential in generating responses that are consistent with clinical guidelines for gout management. The deployment of LLMs in gout patient education and clinical decision support necessitates the simultaneous optimization of both accuracy and readability. Key Points • This study is the first to systematically evaluate the agreement with ACR gout guidelines and readability of three state-of-the-art LLMs (DeepSeek-V3, DeepSeek-R1, GPT-5) in gout management, filling the gap of LLM benchmarking for gout patient education. • Response accuracy was significantly higher for GPT-5 compared to DeepSeek-V3, with no significant difference between GPT-5 and DeepSeek-R1. However, GPT-5 produced texts with the lowest readability. • The study innovatively combined expert-rated accuracy (5-point Likert scale by gout and hyperuricemia specialists) and objective readability metrics (FRE, FKGL, ARI), providing a comprehensive framework for assessing LLM utility in chronic disease self-management. • Findings confirm LLMs' potential for gout patient education but emphasize the need for simultaneous optimization of medical accuracy and health literacy, guiding future LLM refinement for clinical application.
Related Concept Videos
Guidelines for Nursing Documentation I
Factual:
The following points emphasize the significance of upholding accurate and unbiased documentation in healthcare.
Guidelines for Writing Outcome
Patient outcomes reflect the patient's response to the goal rather than what the nurse aims to achieve. Terminology should be observable and measurable to avoid the reader's interpretation. The desired outcome should be realistic and achievable in the designated care timeframe. Expected outcomes should align with adjunctive therapies. The outcome should enhance care...
Urinary Tract Calculi V: Nursing Management
Urinary Tract Calculi III: Medical Management
Drug Dosing in Renal Diseases: Estimation of Glomerular Filtration Rate Based on Serum Creatinine Concentration
Atherosclerosis III: Management
