Related Experiment Video
Updated: Jan 10, 2026

05:48
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
2.0K
MVIB-Lip: Multi-View Information Bottleneck for Visual Speech Recognition via Time Series Modeling
Yuzhe Li1, Haocheng Sun1, Jiayi Cai1
1School of Electronic Engineering, Xi'an University of Posts and Telecommunications, Xi'an 710121, China.
Entropy (Basel, Switzerland)
|November 26, 2025
Summary
This study introduces MVIB-Lip, a novel multi-view learning framework for lipreading (visual speech recognition). It improves accuracy and generalization, especially in low-resource scenarios, by combining landmark trajectories and recurrence plots.
Area of Science:
- Computer Science
- Artificial Intelligence
- Signal Processing
Background:
- Lipreading, or visual speech recognition, interprets speech from lip movements.
- Deep learning methods require extensive data and struggle with generalization.
- Existing approaches often fail in low-resource or speaker-independent settings.
Purpose of the Study:
- To develop a more robust and data-efficient lipreading framework.
- To improve generalization capabilities of visual speech recognition systems.
- To explore multi-view learning for integrating complementary lip movement representations.
Main Methods:
- MVIB-Lip framework integrating raw landmark trajectories (multivariate time series) and recurrence plot (RP) images.
- Transformer encoder for temporal sequences and ResNet-18 for RP image features.
- Fusion of two views using product-of-experts posterior and multi-view information bottleneck.
Main Results:
- MVIB-Lip consistently outperformed handcrafted baselines on OuluVS and a self-collected dataset.
- Demonstrated improved generalization for speaker-independent visual speech recognition.
- Validated the effectiveness of combining recurrence plots with deep multi-view learning.
Conclusions:
- Recurrence plots, combined with deep multi-view learning, offer a principled and data-efficient approach for robust visual speech recognition.
- MVIB-Lip presents a promising direction for overcoming limitations of current lipreading techniques.
- The framework shows potential for real-world applications requiring accurate and generalizable lipreading.
More Related Videos
Related Concept Videos
Chunking and Rehearsal in Sensory Memory
546
Improving short-term memory can be achieved through techniques like chunking and rehearsal. Chunking involves organizing information into larger, more manageable units. This technique is particularly useful for information that exceeds the typical memory span of between five and nine items. For instance, logging into an online account with a password like "ta89vq0179gz" involves grouping letters and numbers into three chunks—ta89, vq01, and 79gz. It makes large amounts of...
546
Classification of Signals
1.3K
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
1.3K
Multi-input and Multi-variable systems
376
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence of...
In the absence of...
376

