Video Experimental Relacionado
Updated: Jan 31, 2026

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Un marco híbrido de CNN y aprendizaje por refuerzo para la identificación del hablante utilizando características de
Fereshteh Manafzadeh Heir1, Hossein Najafzadeh2, Sarvenaz Erfani3
1Department of Computer Engineering, Faculty of Electrical and Computer Engineering, Semnan University, Semnan, Iran.
Abstract:
Speaker identification remains critical in biometric authentication systems, requiring robust feature extraction strategies that capture speaker-specific vocal characteristics. This study introduces a hybrid deep learning architecture integrating Convolutional Neural Networks (CNNs) with Reinforcement Learning (RL) for confidence-aware speaker identification. Two feature extraction methodologies were compared: Method 1 employed Mel-spectrogram representations (80 bins, 20-8000 Hz) with self-attention mechanisms, while Method 2 utilized Continuous Wavelet Transform with Morlet wavelets (128 scales). Both methods were implemented as hybrid CNN-RL architectures and compared against CNN-only baselines. Frameworks were evaluated on LibriSpeech dev-clean dataset (2,703 audio files, 40 speakers) through stratified 5-fold cross-validation. ANOVA assessed discriminative capacity of 22 acoustic features. ANOVA revealed 21 of 22 features demonstrated significant discriminative power (p < 0.05), with entropy exhibiting the strongest effect (F = 39.79, η² = 0.37). Method 1 achieved 87.60% accuracy (95% CI: [83.60%, 91.95%]) and ROC-AUC of 99.54%. Method 2 attained 77.60% accuracy (95% CI: [73.12%, 82.08%]) with ROC-AUC of 98.21%. Ablation studies demonstrated that RL integration provided statistically significant improvements over CNN-only baselines (Method 1: +2.80%, p = 0.0142; Method 2: +3.40%, p = 0.0089) with reduced performance variability. Both architectures demonstrated balanced metrics with Matthews Correlation Coefficients exceeding 0.76. Q-learning integration enabled adaptive decision-making for classification uncertainty, with ablation studies confirming quantitative contributions over conventional CNNs. Results demonstrate that Mel-scale representations with attention mechanisms provide superior discriminative capacity compared to fixed-resolution wavelet transforms for robust biometric authentication.
Videos de Conceptos Relacionados
Hybrid Zones
Continuous -time Fourier Transform
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Corrosion of Reinforcement
However, over time and under certain conditions like carbonation, chloride ingress, and cracking this protective state can be compromised. Steel has areas with...
Reinforcement Schedules
Once a behavior is learned,...
Reinforcements in Concrete

