Related Experiment Video
Updated: Jun 24, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Recurrence plot embeddings as short segment nonlinear features for multimodal speaker identification using air, bone
K Khadar Nawas1, A Shahina2, Keshav Balachandar2
1School of Computer Science and Engineering, Vellore Institute of Technology, Chennai, Tamilnadu, 600127, India.
This study introduces Recurrent Plot (RP) embeddings as novel, nonlinear features for speaker recognition. These embeddings effectively capture unique vocal tract dynamics for accurate speaker identification across different speech transmission modes.
Area of Science:
- Speech processing and acoustics
- Nonlinear dynamics in biological systems
- Machine learning for biometrics
Background:
- Speech production involves a nonlinear Vocal Tract (VT) system.
- Speaker characteristics are often modeled using linear spectral features.
- Nonlinear dynamics of the VT system are underexplored for speaker recognition.
Purpose of the Study:
- To propose Recurrent Plot (RP) embeddings as stand-alone, nonlinear speaker-discriminating features.
- To evaluate the efficacy of RP embeddings for speaker identification across different speech transmission modes (air, bone, throat).
- To assess the performance of unimodal and multimodal systems using RP embeddings.
Main Methods:
- Utilized two datasets: TIMIT speech corpus and a consonant-vowel unimodal syllable dataset.
- Employed Recurrent Plot (RP) embeddings as nonlinear features.
- Conducted closed-set speaker identification experiments using unimodal (Air, Bone, Throat) and multimodal (bimodal, trimodal) systems.
Main Results:
- Unimodal systems trained on RP embeddings achieved high accuracies: Air (95.81%), Bone (98.18%), and Throat (99.74%).
- The best trimodal system (Air-Bone-Throat) reached 99.84% accuracy, comparable to systems using spectrograms and MFCCs.
- A bimodal Bone-Throat system achieved 98.84% accuracy, demonstrating effectiveness without air conduction speech.
Conclusions:
- RP embeddings are significant nonlinear features capable of independent speaker recognition.
- The nonlinear dynamics of the Vocal Tract (VT) system, captured by RP embeddings, are highly speaker-specific.
- This nonlinear feature representation holds potential for advancing both speaker and speech recognition technologies.
Related Concept Videos
Receiver Operating Characteristic Plot
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Residual Plots
When the residual values are plotted against the variable x, it is called a residual...
Determination of Expected Frequency
Relative Frequency Histogram

