Related Experiment Videos
Patient-facing diabetic foot information from large language models: a domain- and source-balanced prompt framework
Yang Wen1, Liyuan Chen2, Jiaping Lan1
1Department of Lower Extremity and Orthopedic Surgery, Suining Central Hospital, Suining, China.
Frontiers in Endocrinology
|August 5, 2026
Summary
Benchmarking large language models (LLMs) for diabetic foot disease information revealed significant differences in response quality. LLM outputs lacked transparency and readability, necessitating clinician oversight for patient decision-making.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Healthcare
- Patient Education
Background:
- Diabetes-related foot disease necessitates prompt identification of risks like neuropathy, ulceration, infection, and ischemia.
- Publicly available large language models (LLMs) offer patient information but lack standardized benchmarking for prompt construction.
- Reproducible methods are needed to evaluate the quality of LLM-generated patient-facing health information.
Purpose of the Study:
- To develop and implement a balanced prompt framework for assessing patient-facing diabetic foot information from LLMs.
- To benchmark the performance of publicly accessible LLMs using a structured, domain- and source-balanced approach.
- To evaluate LLM-generated diabetic foot information under default single-turn, public-interface conditions.
Main Methods:
- A 24-item benchmark prompt set was created, balancing clinical domains with information sources (web trends, Chinese and international guidelines).
- Prompts were submitted to five LLMs (GPT-5.5 Thinking, DeepSeek-V4, Gemini 3.1 Pro, Grok 4.3, Qwen3.6-Max-Preview), generating 120 responses.
- Response quality was evaluated using DISCERN, EQIP, GQS, JAMA criteria for transparency, and readability formulas; a clinical-risk flag was used for safety screening.
Main Results:
- Significant metric-specific variations were observed across the evaluated LLMs.
- Grok 4.3 demonstrated the highest scores for quality and transparency metrics, while DeepSeek-V4 had the lowest readability.
- Visible transparency remained low across all models, and no responses met all readability targets; no critical safety risks were flagged.
Conclusions:
- The developed prompt framework offers a structured method for benchmarking LLM-generated diabetic foot education content.
- Default LLM responses exhibit variable quality, limited transparency, and inadequate readability.
- Patient-facing LLM information for high-risk diabetic foot conditions should not be used independently without professional clinical oversight.