一个研究改进的两阶段双Conv坐标注意力模型,用于声音事件检测和定位
Guorong Chen1, Yuan Yu1, Yuan Qiao1
1School of Intelligent Technology and Engineering, Chongqing University of Science and Technology, No. 20, Daxuecheng East Road, Shapingba District, Chongqing 401331, China.
Sensors (Basel, Switzerland)
|August 29, 2024
概括
本研究引入了一种用于声事件检测和定位 (SELD) 的新模型,该模型通过使用双调节坐标注意模块和增强的反复单位来提高准确性. 在TAU Spatial Sound Events 2019数据集中,TDCAM模型的表现明显优于基线方法.
科学领域:
- 人工智能的人工智能
- 机器学习 机器学习
- 信号处理 信号处理
背景情况:
- 声音事件检测和定位 (SELD) 由于时间和空间的声音事件重叠而具有挑战性.
- 现有的两阶段模型在时间处理限制下扎.
研究的目的:
- 开发一个改进的SELD模型,解决当前方法的局限性.
- 增强SELD.的特征选择和时间建模能力.
主要方法:
- 介绍了以SELD为导向的两阶段双交坐标注意力模型 (TDCAM).
- 集成的双传输坐标注意模块 (DCAM) 用于功能增强.
- 采用了双层双向封闭反复单元 (Bi-GRU) 和数据增强 (频率/时间面具).
主要成果:
- 与基线双阶段网络相比,TDCAM显著提高了SELD性能.
- 废弃性研究证实了DCAM和双层Bi-GRU结构的有效性.
- 该模型展示了对时间特征的增强建模和概括.
结论:
- 拟议的TDCAM模型在SELD.中提供了显著的进步.
- 结合DCAM和增强的循环结构,有效地解决了SELD的挑战.
- 这些发现为未来在多模式事件检测和定位方面的研究提供了坚实的基础.
相关概念视频
Perceiving Loudness, Pitch, and Location
203
The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
203
Difference from Background: Limit of Detection
6.0K
The limit of detection (LOD) is the smallest amount of analyte that can be distinguished from the background noise. The LOD value corresponds to the concentration at which the analyte signal is three times larger than the standard deviation of the blank signal. Below this value, the analyte signal cannot be differentiated from the background noise. It is calculated by dividing the calibration slope by 3 times the standard deviation of the blank signals.
The LOD indicates the presence or absence...
The LOD indicates the presence or absence...
6.0K


