Related Experiment Video
Updated: May 5, 2026

07:24
Using Eye-tracking to Assess the Relative Importance of Visual and Vestibular Input to Subcortical Motion Processing in the Roll Plane
Published on: August 22, 2025
607
A Vision-Based Subtitle Generator: Text Reconstruction via Subtle Vibrations from Videos.
Yan Wang1, Yingchong Wang1, Xiuqi Zhang1
1School of Mechanical Engineering, Beijing Institute of Technology, Haidian District, Beijing 100081, China.
Sensors (Basel, Switzerland)
|March 14, 2026
Summary
This study introduces a Vision-based Subtitle Generator (VSG) that converts sound-induced object vibrations into text. This novel approach uses phase-based motion estimation and a Transformer architecture for accurate speech recovery from visual data.
Area of Science:
- Computer Vision
- Acoustics
- Signal Processing
Background:
- Ambient sound, particularly speech, induces subtle vibrations in everyday objects.
- These vibrations contain acoustic cues that can be potentially decoded into text.
- Applications exist in monitoring and security.
Purpose of the Study:
- To present the Vision-based Subtitle Generator (VSG).
- To enable direct text recovery from high-speed videos of sound-induced object vibrations using a generative approach.
- To reduce the dependency on large volumes of video data for training.
Main Methods:
- Introduced a phase-based motion estimation (PME) technique, treating pixels as "independent microphones" to extract pseudo-acoustic signals.
- Utilized a pretrained Hidden-unit Bidirectional Encoder Representations from Transformers (HuBERT) as the encoder for the VSG-Transformer architecture.
- Leveraged generative approach for vibration-to-text conversion.
Main Results:
- Achieved character error rates of 13.7% (Base) and 12.5% (Large) for text generation from chip bag vibrations.
- Demonstrated the effectiveness of the generative approach in vibration-to-text transcription.
- Showcased robustness to lower sampling rates, maintaining performance with limited temporal sampling.
Conclusions:
- The VSG-Transformer effectively recovers text from sound-induced object vibrations.
- The proposed methods significantly reduce the need for extensive video datasets.
- The system shows promise for real-world applications in diverse acoustic environments.