Related Experiment Video
Updated: Jul 6, 2026

Virtual Agent for Real-Time Motivational Interviewing by Integrating Adaptive Nonverbal Behavior and Language Models
Published on: December 23, 2025
Benchmarking Motivational Interviewing Competence of Large Language Models
Aishwariya Jha1, Prakrithi Shivaprakash1, Lekhansh Shukla2
1Department of Psychiatry, Centre for Addiction Medicine, National Institute of Mental Health and Neuro Sciences (NIMHANS), Bengaluru, India.
Introduction:
Motivational interviewing (MI) promotes behavioural change in substance use disorders. Its fidelity is measured using the Motivational Interviewing Treatment Integrity (MITI) framework. While large language models (LLMs) can potentially generate MI-consistent therapist responses, their competence using MITI is not well-researched, especially in real-world clinical transcripts. We aim to benchmark MI competence of proprietary and open-source models compared to human therapists in real-world transcripts and assess distinguishability from human therapists.
Methods:
We shortlisted 3 proprietary and 7 open-source LLMs from LMArena, evaluated performance using MITI 4.2 framework on two datasets (96 handcrafted model transcripts and 34 real-world clinical transcripts). We generated parallel LLM-therapist utterances iteratively for each transcript while keeping client responses static and ranked performance using a composite ranking system with MITI components and verbosity. We conducted a distinguishability experiment with two independent psychiatrists to identify human-versus-LLM responses.
Results:
All 10 tested LLMs had fair (MITI global scores of >3.5) to good (MITI global scores of >4) competence across MITI measures, and the three best-performing models (gemma-3-27b-it, gemini-2.5-pro, and grok-3) were tested on real-world transcripts. All showed good competence, with LLMs outperforming human-expert in complex reflection percentage (39% vs. 96%) and reflection-question ratio (1.2 vs. >2.8). In the distinguishability experiment, psychiatrists identified LLM responses with only 56% accuracy, with d-prime of 0.17 and 0.25 for gemini-2.5-pro and gemma-3-27b-it, respectively.
Conclusion:
LLMs can achieve good MI proficiency in real-world clinical transcripts using the MITI framework. However, a high complex reflection percentage may result in a technically correct but unnatural conversation style. These findings suggest that even open-source LLMs are viable candidates for expanding MI counselling sessions in low-resource settings, warranting further rigorous clinical validation.
Related Concept Videos
Motivational Bias
Modeling in Therapy
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in situations...
Impression Management Techniques IV: Altercasting
Stereotype Content Model
Strategies of Self-Presentation III: Self-Monitoring
