Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Perceiving Loudness, Pitch, and Location01:21

Perceiving Loudness, Pitch, and Location

405
The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
405
Hearing01:31

Hearing

52.9K
When we hear a sound, our nervous system is detecting sound waves—pressure waves of mechanical energy traveling through a medium. The frequency of the wave is perceived as pitch, while the amplitude is perceived as loudness.
52.9K
Perception of Sound Waves01:01

Perception of Sound Waves

4.6K
The human ear is not equally sensitive to all frequencies in the audible range. It may perceive sound waves with the same pressure but different frequencies as having different loudness. Moreover, the perception of sound waves depends on the health of an individual's ears, which decays with age. The health of one's ears may also be affected by regular exposure to loud noises.
The pitch of a sound depends on the frequency and the pressure amplitude of the source. Two sounds of the same...
4.6K
Auditory Perception01:17

Auditory Perception

561
The auditory system is essential for sound perception, utilizing various critical structures. When sound waves enter the outer ear, they travel through the ear canal and cause the eardrum to vibrate. These vibrations are then transmitted to the middle ear, where three tiny bones – the malleus, incus, and stapes – amplify the sound. This amplification is crucial, as it ensures that the sound vibrations are strong enough to be conveyed to the inner ear. These vibrations then reach the...
561
Aliasing01:18

Aliasing

203
Accurate signal sampling and reconstruction are crucial in various signal-processing applications. A time-domain signal's spectrum can be revealed using its Fourier transform. When this signal is sampled at a specific frequency, it results in multiple scaled replicas of the original spectrum in the frequency domain. The spacing of these replicas is determined by the sampling frequency.
If the sampling frequency is below the Nyquist rate, these replicas overlap, preventing the original...
203

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Development and validation of a competency assessment scale for nurses in reproductive medicine departments.

BMC nursing·2026
Same author

Prevalence of porcine toxoplasmosis using ERP-iELISA in Hubei Province.

Parasitology research·2025
Same author

Bibliometric and visual analysis of global research on endocrine-disrupting chemicals and children's health: evidence, emerging concerns, and research gaps.

Global health action·2025
Same author

Global Research Progress of Mitochondria and Hypoxic-Ischemic Encephalopathy: A Comprehensive Bibliometric Analysis.

Journal of child neurology·2025
Same author

Worldwide research performance on telemedicine for newborns and neonatal intensive care units: A bibliometric and visualization study.

Digital health·2025
Same author

Menstrual disorder is associated with blood type in PCOS patients: evidence from a cross-sectional survey.

BMC endocrine disorders·2025

Related Experiment Video

Updated: Aug 27, 2025

Infant Auditory Processing and Event-related Brain Oscillations
06:34

Infant Auditory Processing and Event-related Brain Oscillations

Published on: July 1, 2015

16.5K

Polyphonic Sound Event Detection Using Temporal-Frequency Attention and Feature Space Attention.

Ye Jin1,2, Mei Wang1,3, Liyan Luo1,2

  • 1Ministry of Education Key Laboratory of Cognitive Radio and Information Processing, Guilin 541006, China.

Sensors (Basel, Switzerland)
|September 23, 2022
PubMed
Summary

This study introduces a new deep learning model, TFFS-CRNN, for classifying complex polyphonic sounds. The model significantly improves sound event detection accuracy by focusing on critical temporal-frequency and feature space information.

Keywords:
convolutional recurrent neural networksfeature aggregationfeature space attentionsound event detectiontemporal-frequency attention

More Related Videos

Mapping Cortical Dynamics Using Simultaneous MEG/EEG and Anatomically-constrained Minimum-norm Estimates: an Auditory Attention Example
08:45

Mapping Cortical Dynamics Using Simultaneous MEG/EEG and Anatomically-constrained Minimum-norm Estimates: an Auditory Attention Example

Published on: October 24, 2012

14.7K
Measuring Attention and Visual Processing Speed by Model-based Analysis of Temporal-order Judgments
13:00

Measuring Attention and Visual Processing Speed by Model-based Analysis of Temporal-order Judgments

Published on: January 23, 2017

10.0K

Related Experiment Videos

Last Updated: Aug 27, 2025

Infant Auditory Processing and Event-related Brain Oscillations
06:34

Infant Auditory Processing and Event-related Brain Oscillations

Published on: July 1, 2015

16.5K
Mapping Cortical Dynamics Using Simultaneous MEG/EEG and Anatomically-constrained Minimum-norm Estimates: an Auditory Attention Example
08:45

Mapping Cortical Dynamics Using Simultaneous MEG/EEG and Anatomically-constrained Minimum-norm Estimates: an Auditory Attention Example

Published on: October 24, 2012

14.7K
Measuring Attention and Visual Processing Speed by Model-based Analysis of Temporal-order Judgments
13:00

Measuring Attention and Visual Processing Speed by Model-based Analysis of Temporal-order Judgments

Published on: January 23, 2017

10.0K

Area of Science:

  • Machine Learning
  • Audio Signal Processing
  • Deep Learning

Background:

  • Polyphonic sound classification is challenging due to discontinuous and unstable time-frequency variations.
  • Traditional methods struggle with characterizing key features in complex sound events, leading to poor performance.
  • Existing models often fail to effectively utilize temporal-frequency and feature space information.

Purpose of the Study:

  • To propose a novel Convolutional Recurrent Neural Network (CRNN) model, TFFS-CRNN, for enhanced polyphonic sound event detection (SED).
  • To improve the characterization of key feature information in polyphonic SED by integrating temporal-frequency and feature space attention mechanisms.
  • To achieve superior classification performance and reduce error rates in polyphonic SED tasks.

Main Methods:

  • Developed a TFFS-CRNN model integrating Log-Mel spectrograms and MFCCs as input features.
  • Incorporated a temporal-frequency (TF) attention module to capture critical TF features.
  • Implemented a feature space (FS) attention module for dynamic weighting of feature dimensions and a Bidirectional Gated Recurrent Unit (BGRU) for contextual learning.

Main Results:

  • The TFFS-CRNN model demonstrated significant improvements on the DCASE 2016 and 2017 Task3 datasets.
  • Achieved a 12.4% and 25.2% higher F1-score compared to winning systems in the DCASE challenges.
  • Reduced the error rate (ER) by 0.41 and 0.37, respectively, indicating enhanced classification accuracy.

Conclusions:

  • The proposed TFFS-CRNN model effectively addresses the challenges of polyphonic sound classification.
  • The integration of TF and FS attention mechanisms enhances feature characterization and model performance.
  • TFFS-CRNN offers superior classification performance and lower error rates for polyphonic sound event detection.