Integrated visual transformer and flash attention for lip-to-speech generation GAN

Qiong Yang1,2, Yuxuan Bai3, Feng Liu1

  • 1School of Computer Science, Xi'an Polytechnic University, Xi'an, 710048, Shaanxi, China.

Scientific Reports
|February 24, 2024
PubMed
Summary

This study introduces FA-GAN, a novel Lip-to-Speech (LTS) model that significantly improves Chinese and English speech recognition accuracy. FA-GAN addresses key challenges in lip-speaking variation and joint movement modeling for better communication.