NaturalL2S: End-to-end high-quality multispeaker lip-to-speech synthesis with differential digital signal processing.

Yifan Liang1, Fangkun Liu1, Andong Li1

  • 1Institute of Acoustics, Chinese Academy of Sciences, Beijing, 100190, China; University of Chinese Academy of Sciences, Beijing, 100049, China.

Summary

This study introduces Natural Lip-to-Speech (NaturalL2S), an end-to-end framework that improves lip-to-speech synthesis by bridging the domain gap in mel-spectrograms. NaturalL2S enhances synthesized speech quality using acoustic inductive priors and a fundamental frequency predictor.