Related Experiment Video
Updated: Aug 6, 2026

Virtual Agent for Real-Time Motivational Interviewing by Integrating Adaptive Nonverbal Behavior and Language Models
Published on: December 23, 2025
Automated Fidelity Monitoring of Lay-Delivered Mental Health Interventions Using Large Language Models: Development
Shadrack Lilan1, Brandon Mochama1, Tom Osborn1
1Shamiri Institute, 13th Floor, CMS Africa, Chania Avenue, Nairobi, 00505, Kenya, +254 (0) 112540760.
Background:
Task-shifting can help close the mental health treatment gap in low- and middle-income countries, but its effectiveness depends on ongoing supervision, which is hard to scale. AI tools that process session recordings and generate structured fidelity feedback could offer a scalable alternative; yet, to our knowledge, none have been developed or validated for lay-delivered, multilingual, group-format interventions in low-resource settings.
Objective:
We developed and pilot-validated shamiriAI (ShamiriAI Institute), an automated fidelity-monitoring tool for lay-delivered mental health interventions, embedded within the Shamiri school-based program in Kenya.
Methods:
Across 6 secondary schools in Ngong Hub, Kajiado County, Kenya (May-September 2025), shamiriAI processed session audio from 47 lay providers through a 5-stage pipeline: ingestion, multilingual automatic speech recognition (ASR) with prosodic feature extraction, personally identifiable information scrubbing, large language model-based fidelity inference, and supervisor reports. The following two aims were assessed: (1) ASR performance on a held-out test set of manually transcribed sessions and (2) interrater reliability between shamiriAI and independent human supervisor ratings across 52 sessions (38 AI-augmented and 14 standard) on 6 domains (Required Contents, Specifics, Thoroughness, Clarity, Skill, and Purity; 1-7 scale). Reliability used intraclass correlation coefficients, Bland-Altman analysis, adjacent-agreement rates, paired t tests with Holm-Bonferroni correction, and Gwet AC2 (ordinal weights) across 3 formulations of the human reference.
Results:
The ASR model achieved a character error rate of 0.19, word error rate of 0.34, and cosine semantic similarity of 0.77, indicating strong meaning preservation in code-switched speech. AI fidelity scores were systematically lower than the human composite overall (mean 5.14, SD 0.77 vs mean 5.93, SD 0.57 ); Δ=-0.79; d=-1.16; P<.001). Primary intraclass correlation coefficients ranged from -0.06 to 0.20 across the 6 domains, and AC2 sensitivity analyses (against each individual rater and the rounded composite) corroborated this dimension-level ordering. Three patterns emerged: large systematic underrating on holistic dimensions (Required Contents: d=-3.48; Clarity d=-1.56); bidirectional medium-effect bias on facilitation dimensions (Thoroughness d=-0.99; Skill d=+0.87); and substantial agreement on Specifics, Skill, and Purity (Gwet AC2 0.69-0.76 against the rounded composite), relative to a human-human AC2 ceiling of 0.42-0.60 estimated in the same dataset. Exploratory, underpowered subgroup and per-arm checks found no preliminary evidence of bias by lay-provider sex or age band or of arm-level differences.
Conclusions:
Within this 52-session pilot, shamiriAI shows technically feasible multilingual ASR and a coherent, dimension-dependent reliability profile. Specifics, Purity, and Skill already reach substantial agreement, while underperformance on holistic dimensions (Required Contents and Clarity) reflects diagnosable misalignments in rubric interpretation and prompt design, specifying a concrete agenda for shamiriAI (version 2; Shamiri Institute). Whether AI-augmented supervision improves provider skill or student mental health outcomes will be tested in a planned cluster-randomized noninferiority trial.