Related Experiment Video
Updated: Mar 13, 2026

A Simplified Technique for Producing an Ischemic Wound Model
Published on: May 2, 2012
Comparative assessment of large language models in diabetic foot infection management: alignment with IWGDF/IDSA
Hongxia Wu1, Jiayi Deng2,3, Xu Qiu2,3
1Emergency Department, Hangzhou Traditional Chinese Medicine Hospital Affiliated to Zhejiang Chinese Medical University, Hangzhou, China.
Objective:
To assess the clinical utility of artificial intelligence (AI) models (ChatGPT-4o, DeepSeek-R1, Grok-3 and Claude-3.7) in aligning with international guidelines for diabetic foot infection (DFI) management.
Background:
AI systems have demonstrated their potential application value in numerous fields. However, the specific effects of these technologies in the medical and health sector still require in-depth exploration. DFI is a relatively common and serious complication among diabetic patients, and the accurate transmission of relevant information is of great significance. Therefore, it is particularly important to evaluate whether artificial intelligence can serve as an effective clinical auxiliary tool.
Methods:
Responses from ChatGPT-4o, DeepSeek-R1, Grok-3 and Claude-3.7 were evaluated against DFI guidelines using four clinical dimensions (Accuracy, Overconclusiveness, Supplementary Value, and Completeness) using a 5-point Likert scale, and assessed for readability using Flesch Reading Ease (FRE) and Flesch-Kincaid Grade Level (FKGL). Statistical analyses included ANOVA and post hoc comparisons.
Results:
No significant differences were found across models for Accuracy and Overconclusiveness (p > 0.05). However, Supplementary Value differed significantly (p < 0.001), the performance of Grok-3 is superior to that of ChatGPT-4o (p < 0.0001), DeepSeek-R1 (p=0.003), and Claude-3.7 (p < 0.0001). Meanwhile, there are significant differences in terms of Completeness (p=0.005), Grok-3 outperforms ChatGPT-4o (p=0.016)and Claude-3.7 (p=0.010) significantly.Readability also varied: DeepSeek-R1 responses were more complex than ChatGPT-4o (p = 0.046).
Conclusion:
All models perform comparably in terms of accuracy and in avoiding over-conclusions. Grok-3 outperformed the other models in the dimensions of complementarity and completeness. DeepSeek-R1 generated the most complex text. These findings validate the feasibility of AI in the standardized management of DFI, but the models still need to be further verified through clinical trials to determine their value in the real-world decision-making process.
More Related Videos
09:15Come to the Light Side: In Vivo Monitoring of Pseudomonas aeruginosa Biofilm Infections in Chronic Wounds in a Diabetic Hairless Murine Model
Published on: October 10, 2017
04:09Prospective, Randomized, and Controlled Study of a Human Umbilical Cord Mesenchymal Stem Cell Injection for Treating Diabetic Foot Ulcers
Published on: March 3, 2023
Related Concept Videos
Diabetes: Management and Pharmacotherapy
Insulin remains the cornerstone of treatment for most patients with type 1 and many...
Peripheral Arterial Disease II: Clinical Manifestations and Diagnostic Evaluation
Peripheral Artery Disease III: Interprofessional Care