在中英双语语境中改进了语音真实性检测
Cheng-Yuan Tsai1, Sheng-Chain Chang2, Chao-Hsiang Hung2
1Forensic Science Division, Ministry of Justice Investigation Bureau, New Taipei City 231, Taiwan.
Sensors (Basel, Switzerland)
|November 9, 2024
概括
这项研究引入了一种新的深度学习模型来检测被改的音频,在多语言环境中显著提高了准确性. 增强的ResNet-LSTM方法在识别复杂的语音伪造攻击方面超越了当前的领先模型.
科学领域:
- 人工智能的人工智能
- 语音处理 语音处理
- 网络安全 网络安全
背景情况:
- 语音技术的普及需要先进的方法来验证语音的真实性.
- 现有的音频改检测系统在不同语言和改类型中往往缺乏稳定性.
- 伪造攻击对基于语音的安全系统构成重大威胁.
研究的目的:
- 开发和评估一个改进的模型来检测被改的音频,特别是解决多语言环境中的挑战.
- 增强深度学习模型的概括能力,以检测音频改.
- 将拟议的模型与最近的竞赛中最先进的方法进行基准测试.
主要方法:
- 开发了一种混合深度学习模型,将增强的ResNet架构与长短期内存 (LSTM) 网络集成在一起.
- 创建了一个双语数据集,将自我录制的中国语音和公共英语音频样本 (VCTK2) 结合起来.
- 使用高级改技术评估模型,如CycleGAN语音转换和自动拼接.
主要成果:
- 拟议的ResNet-LSTM模型在双语数据集上实现了11.62%的相同错误率 (EER) 的优异性能.
- 该模型与ASVSpoof 2021和ADD 2022比赛的领先方法相比表现出了更好的表现.
- 对现实的改场景进行了有效性验证,包括语音转换和音频拼接.
结论:
- 集成的ResNet-LSTM模型为检测被改的音频提供了强大的解决方案,特别是在具有挑战性的多语言环境中.
- 这种方法显著提升了反伪造技术的最新技术.
- 这项工作为更安全的语音通信系统提供了基础.
更多相关视频
08:32Examining Online Syntactic Processing of Spoken Complex Sentences in Chinese Using Dual-Modal Interference Tasks
Published on: September 5, 2019
5.6K
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
403
相关概念视频
Improving Translational Accuracy
2.5K
2.5K
Language and Cognition
325
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
325
