Related Experiment Video
Updated: Aug 27, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Improving ultrasound-based multimodal speech recognition with predictive features from representation learning
Hongcui Wang1, Pierre Roussel2, Bruce Denby2
1Zhejiang University of Water Resources and Electric Power, Hangzhou, China.
Abstract:
Representation learning is believed to produce high-level representations of underlying dynamics in temporal sequences. A three-dimensional convolutional neural network trained to predict future frames in ultrasound tongue and optical lip images creates features for a continuous hidden Markov model based speech recognition system. Predictive tongue features are found to generate lower word error rates than those obtained from an auto-encoder without future frames, or from discrete cosine transforms. Improvement is apparent for the monophone/triphone Gaussian mixture model and deep neural network acoustic models. When tongue and lip modalities are combined, the advantage of the predictive features is reduced.
Related Concept Videos
Improving Translational Accuracy
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Linear Approximation in Frequency Domain
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
Upsampling
Predicting Products: SN1 vs. SN2
With increased substitution on the alkyl halide,...
Nonconscious Mimicry

