Related Experiment Video
Updated: Jul 10, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Speech Perception Improvement Algorithm Based on a Dual-Path Long Short-Term Memory Network
Hyeong Il Koh1, Sungdae Na2, Myoung Nam Kim3
1Department of Medical & Biological Engineering, Graduate School, Kyungpook National University, Daegu 41944, Republic of Korea.
Abstract:
Current deep learning-based speech enhancement methods focus on enhancing the time-frequency representation of the signal. However, conventional methods can lead to speech damage due to resolution mismatch problems that emphasize only specific information in the time or frequency domain. To address these challenges, this paper introduces a speech enhancement model designed with a dual-path structure that identifies key speech characteristics in both the time and time-frequency domains. Specifically, the time path aims to model semantic features hidden in the waveform, while the time-frequency path attempts to compensate for the spectral details via a spectral extension block. These two paths enhance temporal and spectral features via mask functions modeled as LSTM, respectively, offering a comprehensive approach to speech enhancement. Experimental results show that the proposed dual-path LSTM network consistently outperforms conventional single-domain speech enhancement methods in terms of speech quality and intelligibility.
More Related Videos
05:38Interaction between Phonological and Semantic Processes in Visual Word Recognition using Electrophysiology
Published on: June 29, 2021
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Related Concept Videos
Long-term Potentiation
Auditory Pathway
When viewed cross-sectionally, the cochlea reveals the scala vestibuli and scala tympani flanking...