选择性声学特征增强用于语音情感识别与杂的语音
Seong-Gyun Leem1, Daniel Fulford2, Jukka-Pekka Onnela3
1Department of Electrical and Computer Engineering, University of Texas at Dallas, Richardson, TX 75080 USA.
概括
这项研究介绍了一种新的语音增强方法,用于语音情感识别 (SER) 系统. 通过选择性地增强仅仅弱的声学特征,拟议的方法显著提高了在噪音条件下情绪识别性能.
科学领域:
- 语音处理 语音处理
- 机器学习 机器学习
- 声学分析 声学分析
背景情况:
- 现实世界语音情感识别 (SER) 系统面临的挑战是背景噪音.
- 语音增强 (SE) 模块可以提高语音质量,但可能会降低关键的 SER 功能.
- 现有的SE方法有风险改变强大的声学特征,这对于准确的情感识别至关重要.
研究的目的:
- 为在杂环境中运行的SER系统开发一个有针对性的语音增强策略.
- 仅增强弱的声学特征,这些特征会对情绪识别性能产生负面影响.
- 为了保持对环境变化有弹性的强大特征.
主要方法:
- 使用在清洁演讲上训练的多个单一特征声学模型识别弱点.
- 基于性能,强度和联合关节等级的特征排名.
- 选择性增强识别的弱低级描述符 (LLDs),保留强大的特征.
主要成果:
- 直接增强弱LLD的表现优于从完全增强的语音中提取LLD.
- 实现了显著的性能增长:17.7% (激发),21.2% (主导) 和3.3% (价值) 在10dB SNR.
- 超越了在MSP-Podcast集团中增强所有LLD的系统.
结论:
- 针对弱点的有针对性的增强是SER在噪音条件下更有效的策略.
- 拟议的方法保留了对强大的情感识别至关重要的有区别的声学信息.
- 这种方法在各种情感维度中大幅提高了SER准确性.
更多相关视频
相关概念视频
Labeling Emotion
1.0K
Emotional labeling is a cognitive process that involves identifying and naming one's emotions, such as anger, fear, happiness, or sadness. It allows individuals to recognize and express their internal emotional states, a critical aspect of emotional regulation and communication. Labeling emotions requires more than mere recognition; it also involves drawing upon memory and contextual cues to understand the current situation and apply a corresponding emotional label. For instance, feeling...
1.0K
Non-Verbal Cues
784
Non-verbal communication extends beyond gestures and facial expressions to include vocal elements known as paralanguage. Paralanguage consists of non-verbal vocal cues such as pitch, loudness, speech rate, pauses, and non-verbal vocalizations like laughter, sighs, and moans. These elements not only accompany speech but also provide critical emotional and contextual information.The Role of Paralanguage in CommunicationParalanguage adds depth to spoken language by conveying emotions and...
784


