音频伪造检测的混合框架,配有序模型和听觉咬伤袋
Misaj Sharafudeen1, Vinod Chandra S S2, Andrew J3
1Machine Intelligence Research Laboratory, Department of Computer Science, University of Kerala, Thiruvananthapuram, Kerala, India.
Scientific reports
|August 30, 2024
概括
本研究引入了一个新的框架来检测音频DeepFake和逻辑访问攻击在扬声器验证系统. 该方法有效地识别了复杂的AI生成的语音假冒,大大提高了这些威胁的检测率.
科学领域:
- 计算机科学 计算机科学
- 人工智能的人工智能
- 网络安全 网络安全
背景情况:
- 自动扬声器验证系统对于安全至关重要,但容易受到复杂的伪造攻击.
- 逻辑访问攻击和DeepFake音频威胁利用人工智能绕过语音认证.
- 现有的系统很难区分真实的语音和先进的合成音频.
研究的目的:
- 介绍一种用于检测逻辑访问和DeepFake音频伪造攻击的新框架.
- 加强扬声器验证系统对不断变化的安全威胁的稳定性.
- 集成音频功能和时间频率表示,以改进伪造检测.
主要方法:
- 使用顺序预测模型,特别是双向LSTM,从音频数据中提取特征.
- 采用一个新的听觉咬伤袋 (BoAB) 算法来进行特征标准化.
- 应用极端学习机器分类器来区分真实和伪造的语音.
主要成果:
- 在DeepFake音频攻击中实现了1.18%的低等错率 (EER).
- 与最先进的方法相比,DeepFake攻击的EER得到了95.16%的改善.
- 检测到逻辑访问攻击,EER为12.22%,表明挑战级别更高.
结论:
- 拟议的框架在检测DeepFake音频伪造攻击方面表现出很高的有效性.
- 该方法在保护扬声器验证系统免受人工智能驱动的语音操纵方面取得了重大进展.
- 可能需要进一步的研究来提高逻辑访问攻击的检测率.
相关概念视频
Perceiving Loudness, Pitch, and Location
203
The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
203
Classification of Signals
420
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
420
Chunking and Rehearsal in Sensory Memory
186
Improving short-term memory can be achieved through techniques like chunking and rehearsal. Chunking involves organizing information into larger, more manageable units. This technique is particularly useful for information that exceeds the typical memory span of between five and nine items. For instance, logging into an online account with a password like "ta89vq0179gz" involves grouping letters and numbers into three chunks—ta89, vq01, and 79gz. It makes large amounts of...
186
Difference from Background: Limit of Detection
6.0K
The limit of detection (LOD) is the smallest amount of analyte that can be distinguished from the background noise. The LOD value corresponds to the concentration at which the analyte signal is three times larger than the standard deviation of the blank signal. Below this value, the analyte signal cannot be differentiated from the background noise. It is calculated by dividing the calibration slope by 3 times the standard deviation of the blank signals.
The LOD indicates the presence or absence...
The LOD indicates the presence or absence...
6.0K


