Related Experiment Video
Updated: Aug 11, 2025

Modeling Verbal Behavior Deficits with the Stimulus Control Ratio Equation, SCoRE
Published on: May 14, 2019
Towards End-2-end Learning for Predicting Behavior Codes from Spoken Utterances in Psychotherapy Conversations.
Karan Singla1, Zhuohao Chen1, David C Atkins2
1University of Southern California, Los Angeles, USA.
This study introduces a new method for analyzing spoken language without needing transcripts. The approach uses speech features to directly predict behavior codes, achieving results comparable to transcription-based methods.
Area of Science:
- Computational linguistics
- Speech processing
- Behavioral analysis
Background:
- Spoken language understanding typically requires complex pipelines, including Automatic Speech Recognition (ASR).
- Existing methods often depend on generating transcripts before analysis, adding computational steps and potential errors.
- Behavioral coding from speech is a key task in various fields, including psychotherapy.
Purpose of the Study:
- To develop a novel framework for predicting utterance-level labels directly from speech features.
- To eliminate the dependency on Automatic Speech Recognition (ASR) and transcription generation for behavioral coding.
- To enable transcription-free behavioral coding for spoken language analysis.
Main Methods:
- Utilized a pre-trained Speech-2-Vector encoder as a bottleneck to generate word-level speech representations.
- Employed an objective function similar to Word2Vec for the pre-trained encoder to learn speech feature encoding.
- Developed a classifier that uses only speech features and word segmentation information for predicting utterance-level labels.
Main Results:
- The proposed framework achieves competitive performance compared to state-of-the-art methods.
- The model successfully predicts psychotherapy-relevant behavior codes using only speech features.
- Demonstrated the efficacy of transcription-free behavioral coding.
Conclusions:
- The novel framework offers an efficient alternative to traditional ASR-dependent spoken language understanding pipelines.
- Directly predicting labels from speech features simplifies the process and reduces computational overhead.
- This approach holds significant potential for real-time and resource-constrained behavioral analysis applications.
More Related Videos
05:41A Step-by-Step Implementation of DeepBehavior, Deep Learning Toolbox for Automated Behavior Analysis
Published on: February 6, 2020
10:11Portable Intermodal Preferential Looking IPL: Investigating Language Comprehension in Typically Developing Toddlers and Young Children with Autism
Published on: December 14, 2012
Related Concept Videos
Modeling in Therapy
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
Behavior Therapy
Exposure therapy is a cornerstone of behavioral treatment for anxiety disorders. It involves systematic exposure to feared stimuli, either in real...
Psychotherapy
Operant Conditioning Intervention
In operant conditioning, behaviors that are...
Cognitive Therapy
Behaviorism
The core premise of behaviorism is its focus on observable behavior rather than internal thoughts or feelings. This approach argues that true scientific...