Related Experiment Video
Updated: Jun 23, 2025

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
A perceptual similarity space for speech based on self-supervised speech representations.
Bronya R Chernyak1, Ann R Bradlow2, Joseph Keshet1
1Faculty of Electrical & Computer Engineering, Technion-Israel Institute of Technology, Haifa 3200003, Israel.
This study introduces a new method for analyzing speech, using perceptual similarity spaces instead of traditional acoustic features. This approach better explains variations in speech intelligibility, especially for second-language (L2) speakers.
Area of Science:
- Speech processing
- Acoustic phonetics
- Machine learning
Background:
- Speech recognition systems struggle with non-optimal conditions like background noise and second-language (L2) speech.
- Existing methods analyzing spectro-temporal properties of speech leave much intelligibility variation unexplained.
- A need exists for novel approaches to understand speech variability and improve recognition accuracy.
Purpose of the Study:
- To investigate an alternative approach to speech analysis using perceptual similarity spaces.
- To determine if perceptual similarity can explain intelligibility variations in L1 and L2 speech.
- To assess the potential of self-supervised learning for encoding speech distinctions.
Main Methods:
- Utilized self-supervised learning to create a perceptual similarity space for speech samples.
- Encoded speech distinctions without relying on pre-defined acoustic features or speech-to-text alignment.
- Quantified distances between first-language (L1) and L2 English speech samples in this space.
Main Results:
- L2 English speech samples showed greater variability (less tight clustering) in the perceptual space compared to L1 samples.
- Distances within the perceptual similarity space correlated with human listener performance.
- L1 English listeners exhibited lower recognition accuracy for L2 speakers whose speech was more distant in the space.
Conclusions:
- Perceptual similarity, captured by self-supervised learning, offers a powerful new dimension for speech analysis.
- This approach provides a more comprehensive explanation for intelligibility variations than traditional feature-based methods.
- Perceptual similarity spaces may form the foundation for next-generation speech and language analysis tools.
Related Concept Videos
Perception of Sound Waves
The pitch of a sound depends on the frequency and the pressure amplitude of the source. Two sounds of the same...
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Self-Discrepancy Theory
Perceptual Constancy
Size constancy is the recognition that an object remains the same size, even when its image on the retina changes. For instance, a bus is perceived to be large enough to carry people, even if it looks tiny from...
The Representativeness Heuristic
Stereotype Content Model

