A Vision-Based Subtitle Generator: Text Reconstruction via Subtle Vibrations from Videos.

Yan Wang1, Yingchong Wang1, Xiuqi Zhang1

  • 1School of Mechanical Engineering, Beijing Institute of Technology, Haidian District, Beijing 100081, China.

PubMed
Summary

This study introduces a Vision-based Subtitle Generator (VSG) that converts sound-induced object vibrations into text. This novel approach uses phase-based motion estimation and a Transformer architecture for accurate speech recovery from visual data.