Related Experiment Videos
Predicting the perception of performed dynamics in music audio with ensemble learning
Anders Elowsson1, Anders Friberg1
1KTH Royal Institute of Technology, School of Computer Science and Communication, Speech, Music and Hearing, Stockholm, Sweden.
The Journal of the Acoustical Society of America
|April 5, 2017
Summary
Musicians use dynamics to express music, but modeling this is hard. New audio features and machine learning models accurately predict musical dynamics, outperforming human listeners.
Area of Science:
- Music Information Retrieval
- Computational Acoustics
- Machine Learning in Music
Background:
- Musical dynamics are crucial for conveying expression and structure.
- Spectral properties of instruments change with dynamics, yet dedicated audio features are lacking.
- Previous research relied on subjective listener ratings for ground truth dynamics.
Purpose of the Study:
- Develop novel audio features to model musical dynamics.
- Create machine learning models for predicting performed dynamics from audio.
- Evaluate the performance of these models against human perception.
Main Methods:
- Developed new audio features capturing spectral characteristics and sectional spectral flux.
- Collected listener ratings as ground truth for performed dynamics.
- Trained and evaluated three machine learning models, including an ensemble of multilayer perceptrons.
Main Results:
- Achieved a high R-squared value of 0.84 using an ensemble of multilayer perceptrons.
- Model performance approaches the upper bound of accuracy given ground truth uncertainty.
- Outperformed individual human listeners and matched average ratings of multiple listeners.
Conclusions:
- The developed audio features effectively model musical dynamics.
- Machine learning, particularly ensembles, can accurately predict dynamics.
- Source separation is a critical component for effective feature extraction in this domain.