Related Experiment Videos
Decision Support Framework for Quality Assurance and Enhancement of Therapeutic Artificial Intelligence Systems:
Boyoung Kang1, Kyungmin Kwon1, Piao Huilin1
1Department of Applied Artificial Intelligence, Sungkyunkwan University, 25-2, Sungkyunkwan-Ro, Jongno-gu, Seoul, Republic of Korea, 82 1053895996.
JMIR Medical Informatics
|July 23, 2026
Summary
EvaluationPlus enhances therapeutic AI chatbots through expert diagnosis and multi-LLM improvements, significantly boosting quality and user preference for digital mental health tools.
Area of Science:
- Artificial Intelligence in Mental Health
- Digital Health Quality Assurance
- Human-Computer Interaction
Background:
- Therapeutic chatbots are prevalent in digital mental health but lack actionable evaluation pathways for quality improvement.
- Existing evaluation methods are often diagnostic, not translating findings into validated enhancements aligned with healthcare standards.
Purpose of the Study:
- Introduce EvaluationPlus, a decision support framework for a reproducible evaluation-to-enhancement loop in therapeutic AI.
- Demonstrate the feasibility of EvaluationPlus using expert diagnosis, multi-LLM enhancement mapping, and within-subject validation.
Main Methods:
- Iterative enhancement of a GPT 4.0-based chatbot (Dr. CareSam) over 3 cycles.
- Expert clinical psychologists diagnosed deficits using a 7-dimension rubric and think-aloud protocols.
- Three LLMs (GPT 4.0, Claude 4.0 Sonnet, Gemini 2.5 Flash) generated enhancement strategies.
- A participant-blinded, within-subject A/B study (N=15) compared baseline and enhanced chatbot versions.
Main Results:
- The enhanced chatbot showed a 41% increase in therapeutic quality (mean score 5.40 to 7.63).
- Targeted dimensions (active listening, personalization, complex thinking) improved significantly (mean gain +3.04).
- 86.7% of participants preferred the enhanced version; expert evaluation confirmed safety and appropriateness.
Conclusions:
- EvaluationPlus provides a feasible framework for iterative quality assurance of therapeutic AI systems.
- The framework links expert diagnosis with multi-LLM enhancement and validation for reproducible improvements.
- Future research should validate EvaluationPlus with diverse populations and longitudinal outcome assessments.