Related Experiment Video
Updated: Sep 9, 2025

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Interpretable LLM-Based Detection of Loose Associations Using Synthetic Speech Data in Early Psychosis
Enrique Gutiérrez1,2,3, Carlos Quesada1,4, Emily DeFraites3,5,6
1Departamento de Matemática Aplicada a las Tecnologías de la Información y las Comunicaciones, Escuela Técnica Superior de Ingeniería de Sistemas Informáticos, Universidad Politécnica de Madrid, C. de Alan Turing, s/n, Campus Sur, Madrid 28031, Madrid, España.
Background And Hypothesis:
Loose Associations (LA) in speech are key indicators of psychosis risk, notably in schizophrenia. Current detection methods are hampered by subjective evaluation, small samples, and poor generalizability. We hypothesize that combining Large Language Models (LLMs) with machine learning techniques could enhance objective identification of LA through improved semantic and probabilistic linguistic measures.
Study Design:
We propose a novel and reproducible workflow for generating synthetic conversational instances of LA using LLMs, guided by linguistic theory and validated through clinical expert review. This synthetic dataset forms the basis for model training and is complemented by an independently collected dataset for evaluation. Features extracted included traditional clause similarity measures alongside novel surprisal metrics quantifying semantic coherence and unexpected lexical shifts. A parsimonious and interpretable Light Gradient Boosting Machine model was trained using only four features.
Study Results:
The final model achieved high accuracy (83.46%; 95% CI: 82.96-83.95) on the synthetic dataset and robust performance on an independent set (82.36%; 95% CI: 81.94-82.78, AUC: 0.868). Our model outperformed baselines, including similarity-only models and prior thought disorder detection workflows. SHapley Additive exPlanations analysis confirmed the interpretability of the selected features, highlighting semantic coherence and word surprisal as key discriminators.
Conclusions:
Our approach demonstrates that LLM-derived linguistic features substantially enhance the objective, scalable detection of LA. The resulting model achieves high accuracy with minimal complexity, facilitating clinical applicability and interpretability. Future research should integrate additional lexical and contextual dimensions to further refine the identification of thought disorders, ultimately supporting early psychosis intervention.
More Related Videos
05:48Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
08:17A Semantic Priming Event-related Potential ERP Task to Study Lexico-semantic and Visuo-semantic Processing in Autism Spectrum Disorder
Published on: April 12, 2018