构建一个抗性别偏见的超级体,作为语音情感识别的深度学习基线
Babak Abbaschian1, Adel Elmaghraby1
1Computer Science and Engineering Department, University of Louisville, Louisville, KY 40292, USA.
Sensors (Basel, Switzerland)
|April 12, 2025
概括
这项研究引入了一个新的语音情感识别 (SER) 超级集体和深度学习模型,可以提高准确性并减少性别偏见. 数据增强技术进一步提高模型公平性,以获得更好的SER系统.
科学领域:
- 人工智能的人工智能
- 机器学习 机器学习
- 语音处理 语音处理
背景情况:
- 语音情感识别 (SER) 对智能系统至关重要,但在稳定性和偏见方面面临挑战.
- 目前的SER标准已经过时,尽管深度学习架构的进步.
- 现有的SER深度学习模型缺乏关于演讲者的性别和分布之外的数据的彻底检查.
研究的目的:
- 通过合并现有数据库,为SER创建一个全面的超级库.
- 通过使用多种深度学习架构,为SER建立一个新的基准.
- 调查和减轻性别偏见,并改善SER模型中的概括性.
主要方法:
- 通过汇总来自多个SER数据库的数据,构建一个新的超级数据库.
- 与各种深度学习架构进行超级库的基准测试,以设置新的性能基准.
- 实施数据增强策略,以解决和减少固有的数据偏差.
主要成果:
- 与在单个数据集上训练的模型相比,在超级数据集上训练的模型表现出更好的概括性和准确性.
- 在超级语料库上的训练显著减少了SER模型中的性别偏见.
- 数据增强在缓解跨性别和情绪的偏见方面被证明是有效的,有时可以实现完全脱离偏见.
结论:
- 开发的超级库为推进SER研究提供了坚实的基础.
- 在增强型超级体上训练的深度学习模型表现出更好的公平性和性能.
- 数据增强是开发公正和更准确的SER系统的关键技术.
相关概念视频
Stereotype Content Model
13.9K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
13.9K
Cognitive Theories: Schachter-Singer Theory of Emotion
154
Stanley Schachter and Jerome Singer proposed the two-factor theory of emotion, which emphasizes the interplay between physiological arousal and cognitive labeling in forming emotional experiences. This theory suggests that emotions are not simply a result of physiological responses but rather a combination of these responses and the individual's cognitive interpretation of them.
Physiological Arousal and Cognitive Labeling
According to this theory, when an individual experiences...
Physiological Arousal and Cognitive Labeling
According to this theory, when an individual experiences...
154
Force Classification
1.0K
Forces play a crucial role in the study of physics and engineering. They are essential in describing the motion, behavior, and equilibrium of objects in the physical world. Forces can be classified based on their origin, type, and direction of action.
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
1.0K
Confirmation Biases
5.4K
The confirmation bias is the tendency to focus on information that confirms our existing beliefs and ignore information that is inconsistent with our expectations. For example, if you think that your professor is not very nice, you notice all of the instances of rude behavior exhibited by the professor while ignoring the countless pleasant interactions he is involved in on a daily basis. Have you ever fallen prey to the confirmation bias, either as the source or target of such bias?
5.4K
Classification of Signals
342
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
342
Improving Translational Accuracy
8.5K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
8.5K


