Related Experiment Videos
Patient-facing diabetic foot information from large language models: a domain- and source-balanced prompt framework
Yang Wen1, Liyuan Chen2, Jiaping Lan1
1Department of Lower Extremity and Orthopedic Surgery, Suining Central Hospital, Suining, China.
Background:
Diabetes-related foot disease requires timely recognition of neuropathic risk, ulceration, infection, ischemia, offloading needs, and recurrence risk. Publicly accessible large language models (LLMs) may provide patient-facing information, but reproducible prompt construction for benchmarking such outputs remains insufficiently characterized.
Objective:
This study aimed to develop and apply a domain- and source-balanced prompt framework for benchmarking patient-facing diabetic foot information generated by publicly accessible LLMs under default single-turn public-interface conditions.
Methods:
A 24-item benchmark prompt set was generated using a domain- and source-balanced framework incorporating public-query sources and guideline-derived decision-critical content. Six clinical domains were crossed with four source categories: Google Trends, Baidu Zhidao, a PubMed-indexed Chinese diabetic foot guideline, and PubMed-indexed international diabetic foot guidelines. Each prompt was submitted once to GPT-5.5 Thinking, DeepSeek-V4, Gemini 3.1 Pro, Grok 4.3, and Qwen3.6-Max-Preview, yielding 120 responses. Response quality was assessed using DISCERN, EQIP, and GQS; visible transparency-related features were evaluated using JAMA benchmark criteria; readability was assessed using six formulas; and an exploratory potential clinical-risk flag (PCF) screened for overt short-term harm signals. Formal claim-level factual-accuracy review, guideline-concordance adjudication, and hallucination-frequency analysis were not performed.
Results:
Significant metric-specific differences were observed across models. Grok 4.3 recorded the highest observed mean DISCERN, EQIP, GQS, and JAMA-based visible transparency-related scores, whereas DeepSeek-V4 showed the lowest observed mean scores for several readability-grade metrics and the highest mean FRES. Visible transparency-related scores remained low across models. No response was rated as PCF 1 or PCF 2. No response met all predefined readability targets.
Conclusions:
The proposed prompt framework provides a structured basis for public-interface LLM benchmarking in diabetic foot education. Default responses showed metric-specific variation, limited visible transparency, and inadequate readability, and should not be relied upon independently for high-risk diabetic foot decision-making without clinician oversight.