Related Experiment Video
Updated: Oct 22, 2025

Mapping Cortical Dynamics Using Simultaneous MEG/EEG and Anatomically-constrained Minimum-norm Estimates: an Auditory Attention Example
Published on: October 24, 2012
Lip Reading by Alternating between Spatiotemporal and Spatial Convolutions.
Dimitrios Tsourounis1, Dimitris Kastaniotis1, Spiros Fotopoulos1
1Department of Physics, University of Patras, 26504 Rio Patra, Greece.
This study introduces the Alternating Spatiotemporal and Spatial Convolutions (ALSOS) module to improve lip reading (LR) accuracy. By alternating spatial and spatiotemporal convolutions, the ALSOS module enhances feature learning for better speech prediction from visual cues.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Machine Learning
Background:
- Lip reading (LR) involves predicting speech from visual cues.
- Effective feature learning from visual speech sequences is crucial for LR performance.
- Integrating spatiotemporal and spatial information processing is key to enhancing LR systems.
Purpose of the Study:
- To investigate the benefits of alternating spatiotemporal and spatial convolutions for lip reading.
- To introduce and evaluate a novel learnable module, ALSOS, for lip reading.
- To enhance the accuracy of end-to-end trained lip reading systems.
Main Methods:
- A novel module, ALSOS (Alternating Spatiotemporal and Spatial Convolutions), was developed.
- ALSOS integrates 3D (spatiotemporal) and 2D (spatial) convolutions with conversion components for sequence-to-sequence mapping.
- The proposed system incorporates ALSOS within ResNet blocks and uses Temporal Convolutional Networks (TCNs) for classification, trained end-to-end on image sequences.
Main Results:
- The ALSOS module demonstrated improved performance in lip reading tasks.
- Experiments on Greek and English datasets (LRW-500) showed enhanced classification accuracy.
- Integrating ALSOS with ResNet architecture improved performance by capturing temporal information at various spatial scales.
Conclusions:
- The ALSOS module effectively captures spatiotemporal dynamics beneficial for lip reading.
- The proposed ALSOS module offers an advantageous enhancement to existing ResNet-based lip reading systems.
- Alternating spatiotemporal and spatial convolutions provides a valuable approach for improving visual speech recognition.
Related Concept Videos
Convolution Properties II
The width property indicates that if the durations of input signals are T1 and T2, then the width of the output response equals the sum of both durations, irrespective of the shapes of the two functions. For instance, convolving two rectangular pulses with durations of 2 seconds and 1 second results in a function with a width of 3 seconds.
The area property asserts that the area under the...
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Region of Convergence of Laplace Tarnsform
Consider a decaying exponential signal that begins at a specific time. When deriving its Laplace transform, the time-domain variable is replaced with a complex variable. This...

