A standardized naturalistic audio stimulus dataset with unsupervised labeling
Anas Al-Naji1,2, Ricarda I Schubotz3,4, Anoushiravan Zahedi5,6
1Institute of Psychology, University of Muenster, Fliednerstr. 21, 48149, Muenster, Germany. alnaji@uni-muenster.de.
Scientific Data
|August 5, 2026
Summary
This study created a new dataset of naturalistic sounds for cognitive neuroscience research. The audio clips are standardized and rated for emotion and surprise, aiding brain and behavior studies.
Area of Science:
- Cognitive Neuroscience
- Auditory Perception
- Psychology
Background:
- Standardized datasets are crucial for reproducible cognitive neuroscience research.
- Naturalistic auditory stimuli are underutilized in experimental paradigms.
- Existing datasets lack comprehensive emotional and perceptual ratings.
Purpose of the Study:
- To develop a standardized dataset of naturalistic audio stimuli.
- To provide normative ratings for emotional valence, startlingness, and recognizability.
- To enable participant-grounded categorization of auditory objects.
Main Methods:
- Collected 291 diverse audio files, standardized to 1.5s duration.
- Collected ratings from 361 participants on valence, startlingness, and recognizability.
- Utilized unsupervised machine learning for text-based auditory object categorization.
Main Results:
- Audio clips were generally recognizable across participants.
- Significant variation in emotional valence and startlingness ratings observed.
- Derived a data-driven organization of auditory object categories.
Conclusions:
- The dataset offers a valuable resource for cognitive neuroscience and neuroimaging.
- Normative data facilitates the use of naturalistic sounds in experimental designs.
- Machine learning approach successfully organized auditory stimuli based on participant perception.

