Related Experiment Video
Updated: Jun 2, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating the interpretability of clinical speech AI models: Lessons from two user studies
Lingfeng Xu1, Visar Berisha1,2, Julie Liss1
1College of Health Solutions, Arizona State University, 550 N 3rd Street, Phoenix, 85004, Arizona, USA.
Abstract:
The deployment of Artificial intelligence (AI) in clinical speech applications has been limited in large part by the lack of interpretability, which is essential for establishing clinician trust and enabling effective decision support. Although methods such as SHapley Additive exPlanations (SHAP) aim to improve transparency in many clinical domains, their applicability to clinical speech-language pathology practice is uncertain. Since these methods rely on data modalities like acoustic signal features and spectrograms, which are unfamiliar to clinicians and misaligned with clinical workflows, the resulting interpretations may introduce additional burden and bias rather than provide clinically meaningful insight. To better understand this challenge, we conducted two consecutive user studies to systematically evaluate a commonly used SHAP-based interpretation design (a bar chart showing the influence of acoustic features on AI decisions) in dysarthria detection. Building on our prior works, eight factors were examined: faithfulness, computational efficiency, cognitive load, human-AI task performance, mental model, user trust, clinical understandability, and decision relevance. The results reveal a previously unrecognized risk in current interpretation practices. The seemingly intuitive bar-chart design frequently misled participating speech-language pathology (SLP) students to interpret feature influence as an indicator of clinical severity. Other findings include difficulty understanding AI mechanisms, discrepancies between human and model reasoning, and the limited ability of interpretations to address clinical questions. Through this work, we highlight the need for interpretation designs that are more closely aligned with clinical reasoning patterns and suggest practical considerations for developing speech-based AI systems that can be meaningfully integrated into clinical practice.