Related Experiment Video
Updated: Sep 6, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
BPCNN: Bi-Point Input for Convolutional Neural Networks in Speaker Spoofing Detection
1Department of Artificial Intelligence, Kongju National University, Cheonan 31080, Korea.
We introduce bi-point input for convolutional neural networks (CNNs) to process variable-length features like speech. This method improves performance by feeding pairs of segments, enhancing information available to the CNN.
Area of Science:
- Artificial Intelligence
- Machine Learning
- Signal Processing
Background:
- Convolutional Neural Networks (CNNs) typically require fixed-size inputs, posing challenges for variable-length data like speech.
- Existing methods like feature segmentation limit CNNs to processing one segment at a time, hindering comprehensive analysis.
Purpose of the Study:
- To propose a novel input method, "bi-point input," for CNNs to effectively handle variable-length features.
- To enhance the information processing capacity of CNNs by allowing them to consider multiple segments simultaneously.
Main Methods:
- The proposed bi-point input method feeds pairs of segments from a variable-length feature into a CNN concurrently.
- Various combination strategies for these segments are explored, with guidance provided for optimal segment length selection.
- The method was evaluated on spoofing detection tasks using the ASVspoof 2019 database.
Main Results:
- The bi-point input method significantly improved performance in spoofing detection tasks.
- A relative reduction in equal error rate (EER) of approximately 17.2% for logical access (LA) and 43.8% for physical access (PA) was achieved.
Conclusions:
- The bi-point input method offers a more effective way for CNNs to process variable-length features compared to traditional segmentation.
- This approach enhances the contextual information available to the model, leading to improved accuracy in tasks like speaker verification and anti-spoofing.
More Related Videos
Related Concept Videos
Facial Feedback Hypothesis
Nonconscious Mimicry
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
Masking and Demasking Agents
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
Linear Approximation in Frequency Domain
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
¹H NMR: Interpreting Distorted and Overlapping Signals
As Δν decreases and the signals move closer, the doublets appear increasingly distorted. The intensities of the inner lines increase at the cost of those of the outer lines as the signals are...

