Related Experiment Video
Updated: Jun 7, 2026

09:27
Using Eye Movements Recorded in the Visual World Paradigm to Explore the Online Processing of Spoken Language
Published on: October 13, 2018
10.0K
Detecting Forged Audio Files Using "Mixed Paste" Command: A Deep Learning Approach Based on Korean Phonemic Features
1Department of Digital Media, Soongsil University, Seoul 07027, Republic of Korea.
Sensors (Basel, Switzerland)
|March 28, 2024
Summary
This study introduces a novel deep learning model for authenticating voice recordings, achieving 97.5% accuracy in detecting forged audio files. The method effectively identifies manipulated audio evidence, crucial for legal proceedings.
Area of Science:
- Digital Forensics
- Artificial Intelligence
- Audio Engineering
Background:
- Voice recordings are increasingly used as digital evidence in legal cases.
- Allegations of audio file forgery are rising, necessitating robust authentication methods.
- Existing forensic techniques struggle to detect sophisticated audio manipulations like "Mixed Paste".
Purpose of the Study:
- To develop and validate a deep learning methodology for identifying forged voice recordings.
- To specifically address the detection of audio files altered using the "Mixed Paste" technique.
- To enhance the reliability of digital audio evidence in legal contexts.
Main Methods:
- A hybrid deep learning model combining Convolutional Neural Network (CNN) and Long Short-Term Memory (LSTM) was developed.
- Features were extracted from spectrograms and sequences of Korean consonant types for analysis.
- The model was trained on a dataset of authentic and "Mixed Paste" forged audio recordings from iPhones.
Main Results:
- The hybrid deep learning model achieved a high accuracy rate of 97.5% in detecting forged audio files.
- The model's effectiveness was consistent across different smartphone models and audio editing software.
- Validation tests confirmed the model's capability to identify manipulated audio files.
Conclusions:
- The proposed deep learning framework offers a powerful new tool for audio file authentication in digital forensics.
- This research advances audio forensics by introducing a novel hybrid model capable of detecting "Mixed Paste" audio forgeries.
- The findings support the broader application of AI in ensuring the integrity of digital evidence.
Keywords:
Korean phonemic featuresMixed Pasteaudio forgerydeep learningforged smartphone audio filestransition bandMore Related Videos
Related Concept Videos
¹H NMR: Interpreting Distorted and Overlapping Signals
Spin systems where the difference in chemical shifts of the coupled nuclei is greater than ten times J are called first-order spin systems. These nuclei are weakly coupled, and their chemical shifts and coupling constant can generally be estimated from the well-separated signals in the spectrum.
As Δν decreases and the signals move closer, the doublets appear increasingly distorted. The intensities of the inner lines increase at the cost of those of the outer lines as the signals are slanted or...
As Δν decreases and the signals move closer, the doublets appear increasingly distorted. The intensities of the inner lines increase at the cost of those of the outer lines as the signals are slanted or...
Korotkoff Sounds
Korotkoff sounds are the specific sounds heard while measuring blood pressure using a sphygmomanometer, typically with a stethoscope or a Doppler device. They are named after Russian physician Nikolai Korotkov, who first described them in 1905. These sounds correspond to turbulent blood flow in the artery as the blood pressure cuff is gradually released after inflation.
During blood pressure assessment, inflating the cuff 30 millimeters of mercury above the patient's systolic blood pressure...
During blood pressure assessment, inflating the cuff 30 millimeters of mercury above the patient's systolic blood pressure...
Extraction: Advanced Methods
Metal ions can be separated from one another by complexation with organic ligands–the chelating agent– to form uncharged chelates. Here, the chelating agent must contain hydrophobic groups and behave as a weak acid, losing a proton to bind with the metal. Since most organic ligands used in this process are insoluble or undergo oxidation in the aqueous phase, the chelating agent is initially added to the organic phase and extracted into the aqueous phase. The metal-ligand complex is formed in...
Auditory Pathway
Auditory pathways constitute the complex neural circuits responsible for transmitting and interpreting auditory information from the peripheral auditory system to the brain. Sound waves are initially captured by the outer ear, funneled through the ear canal, and reach the tympanic membrane (eardrum). These vibrations are transmitted via the middle ear's ossicles to the inner ear's cochlea.
When viewed cross-sectionally, the cochlea reveals the scala vestibuli and scala tympani flanking the...
When viewed cross-sectionally, the cochlea reveals the scala vestibuli and scala tympani flanking the...
Air-entraining Agents
Air-entraining agents improve the durability and workability of concrete in climates with frequent freezing and thawing. These agents prevent cracks by introducing small air bubbles into the mix, creating spaces accommodating water expansion when temperatures drop. The air-entraining agents lower the surface tension of water, forming stable, small air bubbles. This method is more effective than having accidental large voids, as the intentional, smaller, and evenly distributed air voids improve...
Auditory Perception
The auditory system is essential for sound perception, utilizing various critical structures. When sound waves enter the outer ear, they travel through the ear canal and cause the eardrum to vibrate. These vibrations are then transmitted to the middle ear, where three tiny bones – the malleus, incus, and stapes – amplify the sound. This amplification is crucial, as it ensures that the sound vibrations are strong enough to be conveyed to the inner ear. These vibrations then reach the cochlea, a...

