UrduSER: 乌尔都语语言语音情感识别的综合数据集
Muhammad Zaheer Akhtar1,2, Rashid Jahangir2, QuratUl Ain1
1Department of Information Technology, The Islamia University of Bahawalpur, Bahawalpur 63100, Pakistan.
Data in brief
|June 11, 2025
概括
创建了一个新的乌尔都语音情感识别数据集 (UrduSER),以解决对分析乌尔都语音情感的资源缺乏的问题. 这个数据集包含了来自专业参与者的多样化,现实世界的对话,增强了乌尔都语SER的机器学习模型.
科学领域:
- 语音情感识别 (SER) 是一种语言识别技术.
- 计算语言学 计算语言学
- 机器学习 机器学习
背景情况:
- 语音情感识别 (SER) 对于理解人机交互至关重要,并具有重要的社会文化和商业应用.
- 乌尔都语的现有数据集,第10个最常说的语言,在范围,情感范围和对话多样性上是有限的,阻碍了现实世界的SER应用.
- 乌尔都语SER存在重大研究缺口,原因是缺乏全面和多样化的语音情感数据集.
研究的目的:
- 为乌尔都语语言 (UrduSER) 开发一个全面和平衡的语音情感识别数据集.
- 通过结合多样化,现实世界的对话和更广泛的情感来解决现有的乌尔都语SER数据集的局限性.
- 促进机器学习和深度学习模型在乌尔都语语音情感分析方面的进步.
主要方法:
- 收集了来自10位巴基斯坦职业演员的3500个语音信号 (性别和年龄平衡),这些语音信号来自于YouTube剧集和电视剧.
- 包括七种不同的情绪状态:愤怒,恐惧,无聊,厌恶,快乐,中立和悲伤,每种情绪有500个样本.
- 确保了对话多样性,每个发言都有独特的内容,并为每个音频样本提供了详细的元数据,包括脚本.
主要成果:
- 开发了UrduSER数据集,这是一个全面的资源,包含3500种不同的语音信号.
- 数据集包括一个平衡的分布,每情绪500个样本和每情绪50个样本.
- 专家验证证实了数据集的有效性,可靠性和适用于研发的适用性.
结论:
- 乌尔都SER数据集有效地填补了乌尔都语语音情感识别的关键研究缺口.
- 它的多样化,现实世界的性质和全面的元数据增强了其用于训练强大的SER模型的实用性.
- 预计这项资源将大大促进乌尔都语SER的研发,使在口语中更准确,更细微的情感检测成为可能.
更多相关视频
05:51Exploring the Use of Isolated Expressions and Film Clips to Evaluate Emotion Recognition by People with Traumatic Brain Injury
Published on: May 15, 2016
9.0K
07:12Protocol for Data Collection and Analysis Applied to Automated Facial Expression Analysis Technology and Temporal Analysis for Sensory Evaluation
Published on: August 26, 2016
9.4K
相关概念视频
Classification of Signals
418
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
418
Heart Sounds
1.9K
Heart sounds are generated by the turbulence in blood flow due to the closing of heart valves. These sounds are best perceived slightly away from the valves, where the blood flow disseminates the sound.
Auscultation is the process of listening to these internal body sounds using a stethoscope. The heart produces four types of sounds, but only two—S1 and S2—can usually be heard with a stethoscope.
S1, also known as the "lub" sound, is caused by the closure of atrioventricular (A-V)...
Auscultation is the process of listening to these internal body sounds using a stethoscope. The heart produces four types of sounds, but only two—S1 and S2—can usually be heard with a stethoscope.
S1, also known as the "lub" sound, is caused by the closure of atrioventricular (A-V)...
1.9K
