Related Experiment Video
Updated: Jun 12, 2026

An Automated System for Sound Localization Testing in Hearing-Impaired Listeners
Published on: March 13, 2026
Sound retrieval and ranking using sparse auditory representations
Richard F Lyon1, Martin Rehn, Samy Bengio
1Google, Mountain View, CA 94043, USA. dicklyon@google.com
New auditory models significantly outperform traditional methods for sound recognition. These advanced models improve sound classification accuracy by 18% in large-scale tests, enhancing machine understanding of everyday sounds.
Area of Science:
- Acoustics and Signal Processing
- Machine Learning for Audio Analysis
- Computational Auditory Scene Analysis
Background:
- Effective sound representation is crucial for systems understanding human auditory environments.
- Evaluating sound representations requires large-scale, quantitative frameworks.
- Machine vision techniques can be adapted for audio processing tasks.
Purpose of the Study:
- To quantitatively evaluate different sound representations in a large-scale sound-ranking task.
- To compare novel auditory models against conventional Mel-frequency cepstral coefficients (MFCCs).
- To adapt and apply the passive-aggressive model for image retrieval (PAMIR) to audio feature extraction.
Main Methods:
- Utilized a sound-ranking framework adapted from the passive-aggressive model for image retrieval (PAMIR).
- Compared adaptive pole-zero filter cascade (PZFC) auditory filter banks with sparse-code feature extraction.
- Evaluated stabilized auditory images with multiple vector quantizers against conventional MFCC front ends.
Main Results:
- Auditory models demonstrated a significant advantage over vector-quantized MFCCs.
- The best auditory model achieved 73% precision at top-1 and 35% average precision.
- This represents an 18% improvement compared to the best-performing MFCC front end.
Conclusions:
- Advanced auditory models offer superior performance for large-scale sound recognition tasks.
- Sparse-code feature extraction from auditory images provides a more discriminative representation than MFCCs.
- The PAMIR framework is effective for evaluating and developing robust auditory feature representations.
More Related Videos
05:48Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis
Published on: August 9, 2024
11:39Assessment of Audio-Tactile Sensory Substitution Training in Participants with Profound Deafness Using the Event-Related Potential Technique
Published on: September 7, 2022
Related Concept Videos
Sound Intensity Level
The human ear can perceive an extensive range of sound intensity, necessitating the use of the logarithmic scale to define a physical quantity—the intensity level. It is a ratio of two intensities and hence a...
Sound Intensity
Chunking and Rehearsal in Sensory Memory
Perceiving Loudness, Pitch, and Location
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by identifying...
Auditory Pathway
When viewed cross-sectionally, the cochlea reveals the scala vestibuli and scala tympani flanking the...
Auditory Perception