Related Experiment Videos
ARTSTREAM: a neural network model of auditory scene analysis and source segregation
Stephen Grossberg1, Krishna K Govindarajan, Lonce L Wyse
1Department of Cognitive and Neural Systems, Center for Adaptive Systems, Boston University, 677 Beacon Street, Boston, MA 02215, USA. steve@bu.edu
Summary
The ARTSTREAM model explains auditory scene analysis, enabling the brain to separate overlapping sounds using pitch and spatial cues. This neural network simulates how we solve the cocktail party problem by grouping sound frequencies into distinct streams.
Area of Science:
- Auditory neuroscience
- Computational auditory scene analysis
- Neural network modeling
Background:
- The brain must segregate complex auditory scenes with overlapping sound sources and noise.
- Auditory scene analysis allows for the perception of distinct sound streams, akin to solving the cocktail party problem.
Purpose of the Study:
- To present the ARTSTREAM neural network model for auditory scene analysis.
- To elucidate the mechanisms by which the brain groups spectral components into coherent auditory streams based on pitch and spatial cues.
- To explain how multiple streams are distinguished and separated.
Main Methods:
- The ARTSTREAM model transforms sound into frequency-specific activations across a spectral stream layer.
- It utilizes bottom-up spectral-pitch resonances formed by mutual reinforcement between spectral representations and pitch expectations.
- The model incorporates Adaptive Resonance Theory (ART) principles, including harmonic filtering, top-down expectations, and spectral component suppression.
Main Results:
- The model demonstrates how spectral-pitch resonances enable coherent stream formation and tracking in noisy environments.
- It simulates auditory grouping phenomena, such as the bounce percept, and illusory percepts like auditory continuity and scale illusions.
- Spatial location cues are shown to disambiguate sources with similar spectral characteristics.
Conclusions:
- The ARTSTREAM model provides a framework for understanding auditory stream segregation through resonance mechanisms.
- The findings support the hypothesis that ART-like mechanisms are fundamental to auditory processing at multiple levels.
- The model's success in simulating psychophysical data suggests its potential for explaining complex auditory perception, including speech perception.