Related Experiment Video
Updated: Mar 24, 2026

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
The Test of Auditory-Vocal Affect (TAVA) dataset.
Juan S Gómez-Cañón1,2, Camille Noufi2,3, Jonathan Berger2
1Stanford University School of Medicine, Department of Psychiatry and Behavioral Sciences, Stanford, CA, USA.
This study introduces the TAVA dataset, featuring audio and transformed electroglottographic (tEGG) speech signals. It aids Speech Emotion Recognition (SER) research by separating vocal affect from linguistic content.
Area of Science:
- Linguistics
- Psychology
- Computer Science
Background:
- Speech Emotion Recognition (SER) research often struggles to disentangle linguistic content from affective vocal cues.
- Existing datasets may not adequately isolate paralinguistic information for focused analysis.
- Vocal affect plays a crucial role in communication, particularly for individuals with language impairments.
Purpose of the Study:
- To introduce the TAVA dataset, a novel resource for Speech Emotion Recognition (SER).
- To facilitate research by disentangling paralinguistic (affective) and linguistic information in speech.
- To enable studies on affect perception in clinical and non-clinical populations.
Main Methods:
- Developed the TAVA dataset comprising 352 audio recordings of emotionally expressive English speech.
- Created transformed electroglottographic (tEGG) versions of speech signals to suppress phonetic content while preserving affective cues.
- Collected over 120,000 crowd-sourced ratings for valence, arousal, and dominance on both original and tEGG signals.
Main Results:
- The TAVA dataset provides paired audio and tEGG signals, allowing for direct comparison of affect perception.
- Crowd-sourced ratings offer detailed insights into how affective cues are perceived across different signal types.
- tEGG signals effectively isolate vocal affect, demonstrating their utility in SER research.
Conclusions:
- The TAVA dataset is a valuable resource for advancing SER by offering phoneme-reduced, affect-rich speech representations.
- This dataset can support research into vocal affect sensitivity in populations with communication difficulties.
- It enables dissociation of linguistic and affective processing, broadening SER applications.
More Related Videos
07:52Author Spotlight: Investigating Vocal Information Representation in Small Primates and Its Alteration by Psychiatric Disorders Using Noninvasive EEG
Published on: July 26, 2024
06:22Machine Learning-Based Cough Tone Classification: Diagnostic Exploration of Chronic Obstructive Pulmonary Disease and Respiratory Tract Infections
Published on: September 19, 2025
Related Concept Videos
Auditory Perception
Auditory Pathway
When viewed cross-sectionally, the cochlea reveals the scala vestibuli and scala tympani flanking...
Facial Feedback Hypothesis
Non-Verbal Cues
Labeling Emotion
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...