Related Experiment Video
Updated: May 11, 2026

05:48
Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis
Published on: August 9, 2024
An episodic memory-based solution for the acoustic-to-articulatory inversion problem
1Université de Lorraine, Laboratoire Lorrain de Recherche en Informatique et ses Applications, Unité de Recherche Mixte 7503, Vandœuvre-lès-Nancy, F-54506, France.
The Journal of the Acoustical Society of America
|May 10, 2013
Summary
This study introduces a generative episodic memory (G-Mem) for acoustic-to-articulatory inversion. G-Mem effectively models articulatory dynamics and generalizes beyond training data, achieving high accuracy.
Area of Science:
- Speech processing
- Articulatory phonetics
- Machine learning
Background:
- Acoustic-to-articulatory inversion aims to reconstruct speech articulation from audio signals.
- Episodic memory models offer a data-driven approach, avoiding assumptions about the acoustic-articulatory mapping.
- Existing episodic memory models struggle with limited training data, hindering their inversion capabilities.
Purpose of the Study:
- To propose a novel generative episodic memory (G-Mem) for acoustic-to-articulatory inversion.
- To enhance the generalization capabilities of episodic memory models for speech articulation.
- To evaluate G-Mem's performance against established methods using real-world articulatory data.
Main Methods:
- Developed a generative episodic memory (G-Mem) capable of producing novel articulatory trajectories.
- Utilized electromagnetic articulography (EMA) corpora for English and French speech.
- Compared G-Mem with codebook-based and concatenative episodic memory methods.
Main Results:
- G-Mem achieved a root-mean-square error of 1.65 mm and a correlation of 0.71.
- The proposed method demonstrated effective modeling of articulatory dynamics.
- G-Mem showed strong generalization capabilities on unseen data.
Conclusions:
- Generative episodic memory (G-Mem) provides an effective solution for acoustic-to-articulatory inversion.
- G-Mem outperforms traditional episodic memory and codebook methods in accuracy and generalization.
- This approach holds promise for advancing speech production research and applications.
Related Concept Videos
Integration by Parts: Problem Solving
Smart speakers process voice commands by modeling audio inputs as piecewise functions and analyzing them through integration against trigonometric functions, such as cosine. This mathematical approach is fundamental in signal processing, where complex sound waves are decomposed into simpler frequency components.Consider a definite integral involving a piecewise function multiplied by a cosine function. Because the function is defined differently over separate intervals, the integral is split...
Chunking and Rehearsal in Sensory Memory
Improving short-term memory can be achieved through techniques like chunking and rehearsal. Chunking involves organizing information into larger, more manageable units. This technique is particularly useful for information that exceeds the typical memory span of between five and nine items. For instance, logging into an online account with a password like "ta89vq0179gz" involves grouping letters and numbers into three chunks—ta89, vq01, and 79gz. It makes large amounts of information more...

