Related Experiment Video
Updated: Sep 2, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Benchmarking publicly accessible large language models for English-language patient-facing acute pancreatitis
Biao Jiang1, Hongxi Sun2, Linlin Chen1
1Ward I, Department of Gastroenterology, Suining Central Hospital, Suining, Sichuan, China.
Background:
Patients increasingly rely on large language models (LLMs) for health information, yet their suitability for decision-critical conditions such as acute pancreatitis remains unclear. Given that acute pancreatitis requires timely symptom recognition, severity assessment, treatment decision-making, recurrence prevention, and follow-up management, LLM-generated information should demonstrate reliability, transparency, and readability.
Objectives:
To evaluate the informational quality, visible transparency-related features, and readability of English-language responses generated by five publicly accessible LLMs to standardized patient-facing questions on acute pancreatitis.
Methods:
This cross-sectional benchmark study developed 24 English-language, single-intent questions on acute pancreatitis across six clinical domains using public search intents and guideline-derived decision-critical content. The analysis focused exclusively on English-language patient-facing responses. Each question was submitted once to GPT-5.4 Thinking, DeepSeek-V3.2, Gemini 3.1 Pro, Grok 4.3, and Qwen3.6-Max-Preview, generating 120 responses. Anonymized responses were independently evaluated by two blinded gastroenterologists using DISCERN, Ensuring Quality Information for Patients (EQIP), Global Quality Scale (GQS), and Journal of the American Medical Association (JAMA) benchmark criteria. Readability was assessed using six established formulas. A structured response-level safety analysis was added to evaluate factual inaccuracies, clinical hallucinations, and clinical safety-risk severity. Between-model differences were analyzed with Friedman tests followed by post hoc paired Wilcoxon signed-rank tests with Holm correction.
Results:
Significant between-model differences were observed across all quality, transparency-related, and readability outcomes. Grok 4.3 achieved the highest mean DISCERN, GQS, EQIP, and JAMA scores, reflecting the strongest informational quality profile according to the predefined quality instruments, although visible transparency cues remained limited across all models. In the added safety analysis, factual inaccuracies and clinical hallucinations were each identified in 16 of 120 responses (13.3%), whereas clinical safety-risk signals were identified in 7 responses (5.8%), all of which were adjudicated as score 1 (low risk) under the predefined clinician-rated rubric; no moderate- or high-risk clinical safety event was adjudicated. DeepSeek-V3.2 demonstrated the most favorable readability profile, with the highest Flesch Reading Ease score (44.42 ± 12.27), which nevertheless remained substantially below the recommended threshold of ≥ 80. None of the 120 responses satisfied all six predefined readability targets. All model-level readability distributions differed significantly from recommended thresholds in the direction of poorer readability.
Conclusion:
Publicly accessible LLMs generated English-language responses on acute pancreatitis with variable informational quality, limited visible transparency cues, and consistently inadequate readability under standardized default public-interface conditions. Because the primary analysis used a single-generation cross-sectional design, model rankings should be interpreted as performance snapshots rather than definitive or temporally stable hierarchies. Higher quality scores did not necessarily establish factual accuracy, clinical safety, patient-education-level readability, or stronger response-level transparency. Current public-interface LLM outputs may serve as clinician-reviewed drafts for patient education but should not function as standalone patient resources. Future AI-based health information systems should strengthen clinical completeness, plain-language communication, visible evidence support, actionability, and clinician-supervised safeguards.
Related Concept Videos
Acute Pancreatitis II: Clinical Manifestations and Management
Chronic Pancreatitis II: Collaborative Care
Assessment:
Acute Pancreatitis I: Introduction
