Related Experiment Video
Updated: May 19, 2026

Virtual Agent for Real-Time Motivational Interviewing by Integrating Adaptive Nonverbal Behavior and Language Models
Published on: December 23, 2025
Evaluating the Accuracy, Quality, and Reproducibility of AI Chatbot-Generated Diabetes Self-Management Education and
Teng-Hung Yu1,2, Hui-Chun Hsu3, Yau-Jiunn Lee4
1Division of Cardiology, Department of Internal Medicine, E-Da Hospital, I-Shou University, Kaohsiung.
Background:
The rapid development of artificial intelligence, particularly large language models (LLMs) such as ChatGPT, Gemini, and Claude, offers new opportunities to scale and personalize diabetes self-management education and support (DSMES). This study evaluated the quality, accuracy, and reproducibility of DSMES plans generated by ChatGPT-5 using the ADCES7 Self-Care Behaviors™ framework.
Methods:
Eleven virtual patient profiles representing five key DSMES time points (diagnosis, annual review, complicating factors, transitions, and ongoing care) were analyzed using structured prompts with retrieval-augmented generation. Generated plans were assessed using a DSMES Process Evaluation Checklist integrating IMPACTS, QAMAI, PDQI-9, and Donabedian's frameworks across ten domains. Content validity was confirmed by five experts (S-CVI: 0.982 for appropriateness; 0.964 for clarity). Ten certified diabetes educators (mean age 47.9 years; mean experience 14 years) independently scored the plans. Internal consistency, inter-rater reliability, and reproducibility were examined.
Results:
Mean domain scores ranged from 29.0 ± 2.0 (patient engagement; 58.0%) to 42.7 ± 2.0 (cultural and linguistic appropriateness; 85.5%). ChatGPT-5 demonstrated strengths in clarity (85.1%) and goal-directedness (81.1%), generating evidence-based plans within three to five minutes, but showed limitations in patient engagement. Moderate performance was observed for accuracy (74.9%), relevance (78.7%), individualization (79.6%), and feasibility (75.3%). Reliability was high (Cronbach's α = 0.889; intra-class correlation coefficient = 0.837). Reproducibility was moderate, with intra-assay coefficients of 0.04 to 0.15 and inter-assay coefficients of 0.10 to 0.27.
Conclusions:
The DSMES process evaluation checklist is valid and reliable. ChatGPT-5 shows promise as a scalable DSMES decision-support tool, though improvements in personalization, patient engagement, and reproducibility are needed before broader implementation.
Related Concept Videos
Errors occurring during blood pressure monitoring
Several factors...
Nursing Process for Patient and Caregiver Teaching I: Assessment and Diagnosis
It is critical to determine the patient's learning needs during the assessment. Determination of learning needs compounds data from the...