Related Experiment Video
Updated: Jul 7, 2026

07:21
Automated Interactive Video Playback for Studies of Animal Communication
Published on: February 9, 2011
An HMM-based speech-to-video synthesizer
J J Williams1, A K Katsaggelos
1Dept. of Electr. and Comput. Eng., Northwestern Univ., Evanston, IL, USA.
IEEE Transactions on Neural Networks
|February 5, 2008
Summary
This study introduces a novel correlation hidden Markov model (HMM) for synthesizing visual speech from audio. This approach improves speechreading accuracy and reduces training data requirements.
Area of Science:
- Speech processing
- Computer vision
- Multimedia communication
Background:
- Broadband communication systems enable multimedia telephony, integrating visual information with audio.
- Synthesizing visual articulatory movements from acoustic speech signals is crucial for enhancing speechreading capabilities.
- Existing narrowband systems for speech can be leveraged for visual speech synthesis.
Purpose of the Study:
- To develop a hidden Markov model (HMM)-based visual speech synthesizer for generating articulatory movements from acoustic speech signals.
- To introduce a novel correlation HMM that integrates independently trained acoustic and visual HMMs for speech-to-visual synthesis.
- To enhance flexibility in model topology selection and reduce training data needs compared to earlier integration methods.
Main Methods:
- Decomposition of the modeling task into key stages for HMM application.
- Judicious determination of observation vector components for each stage.
- Development and application of a novel correlation HMM for integrating acoustic and visual HMMs.
Main Results:
- Objective experiments showed a 37.4% reduction in time alignment errors compared to conventional temporal scaling.
- Subjective evaluations indicated an increase in speech understanding using the proposed model.
- The correlation HMM allows for greater flexibility in acoustic and visual HMM topologies.
Conclusions:
- The proposed correlation HMM effectively synthesizes visual speech from acoustic signals.
- This method offers improved accuracy and efficiency in visual speech synthesis.
- The approach has the potential to significantly enhance multimedia communication systems.
Related Concept Videos
Synthetic Disvision of Polynomials
Synthetic division is an efficient algorithmic approach for dividing a polynomial by a linear binomial of the form x - c, where c is a real number. This method is helpful due to its streamlined process, which avoids the more cumbersome steps involved in the traditional long division of polynomials. It simplifies computation and serves as a practical tool for evaluating polynomials and identifying their factors.To perform synthetic division, one begins by listing the coefficients of the...
Downsampling
When considering a sampled sequence with zero values between sampling instants, one can replace it by taking every N-th value of the sequence. At these integer multiples of N, the original and sampled sequences coincide. This process, known as decimation, involves extracting every N-th sample from a sequence, thereby creating a more efficient sequence.
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...
Amplifying Signals via Enzymatic Cascade
When a ligand binds to a cell-surface receptor, the receptor's intracellular domain changes shape, which may either activate its enzyme function or allow its binding to other molecules. The initial signal is amplified by most signal transduction pathways. This means that a single ligand molecule can activate multiple molecules of a downstream target. Proteins that relay a signal are most commonly phosphorylated at one or more sites, activating or inactivating the protein. Kinases catalyze the...
Phasor Arithmetics
Phasors and their corresponding sinusoids are interrelated, offering unique insights into the behavior of alternating current (AC) circuits. One way to understand this relationship is through the operations of differentiation and integration in both the time and phasor domains.
When the derivative of a sinusoid is taken in the time domain, it transforms into its corresponding phasor multiplied by j-omega (jω) in the phasor domain, where j is the imaginary unit, and ω is the angular frequency.
When the derivative of a sinusoid is taken in the time domain, it transforms into its corresponding phasor multiplied by j-omega (jω) in the phasor domain, where j is the imaginary unit, and ω is the angular frequency.
Reconstruction of Signal using Interpolation
Signal processing techniques are essential for accurately converting continuous signals to digital formats and vice versa. When a continuous signal is sampled with a period T, the resulting sampled signal exhibits replicas of the original spectrum in the frequency domain, spaced at intervals equal to the sampling frequency. To handle this sampled signal, a zero-order hold method can be applied, which creates a piecewise constant signal by retaining each sample's value until the next sampling...
Buffer Effectiveness
Buffer solutions do not have an unlimited capacity to keep the pH relatively constant . Instead, the ability of a buffer solution to resist changes in pH relies on the presence of appreciable amounts of its conjugate weak acid-base pair. When enough strong acid or base is added to substantially lower the concentration of either member of the buffer pair, the buffering action within the solution is compromised.
The buffer capacity is the amount of acid or base that can be added to a given volume...
The buffer capacity is the amount of acid or base that can be added to a given volume...