Related Experiment Video
Updated: Sep 16, 2025

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
MsDUNE: A multi-scale masked temporal fusion framework for speaker-independent lipreading via Dirichlet uncertainty
Jinghan Wu1, Xingwei An2, Yakun Zhang3
1Academy of Medical Engineering and Translational Medicine, Tianjin University, Tianjin, 300072, China; Tianjin Artificial Intelligence Innovation Center, Tianjin, 300450, China.
None:
Lipreading, the task of recognizing speech based on visual cues from lip movements, typically requires a substantial amount of labeled training data to achieve optimal performance. However, this task is highly sensitive to variations among speakers, often resulting in significantly degraded recognition accuracy for unseen speakers. In this work, we introduce a novel framework, multi-scale masked temporal fusion with Dirichlet uncertainty estimation (MsDUNE), designed to mitigate the feature distribution disparities across different speakers. The proposed framework leverages a Dirichlet distribution to parameterize the latent space of a single feature branch, which is then quantitatively assessed through evidence and belief masses. Furthermore, MsDUNE calibrates multi-scale feature distributions by accounting for the mutual influence of feature beliefs between two branches, thereby enhancing the generalization capability of the lipreading model. We validate our approach through extensive experiments conducted on two widely recognized benchmarks, LRW-ID and AV Letters, as well as a self-collected lipreading dataset, CVSR100. The experimental results highlight the state-of-the-art performance of our method, particularly in scenarios involving unseen or overlapping speakers.
Related Concept Videos
Masking and Demasking Agents
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
Uncertainty: Overview
Facial Feedback Hypothesis
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...

