Related Experiment Video
Updated: Feb 4, 2026

Ultrasound Images of the Tongue: A Tutorial for Assessment and Remediation of Speech Sound Errors
Published on: January 3, 2017
Multi-resolution speech analysis for automatic speech recognition using deep neural networks: Experiments on TIMIT.
Doroteo T Toledano1, María Pilar Fernández-Gallego1, Alicia Lozano-Diez1
1AuDIaS - Audio, Data Intelligence and Speech, Universidad Autónoma de Madrid, Madrid, Spain.
Multi-resolution speech analysis, exploring different time-frequency resolutions, shows promise for improving automatic speech recognition (ASR) systems using deep neural networks (DNNs). While generally beneficial, combining multi-resolution with advanced features yielded only modest gains.
Area of Science:
- Speech Processing
- Acoustic Modeling
- Machine Learning for Speech Recognition
Background:
- Traditional Automatic Speech Recognition (ASR) relies on Mel-Frequency Cepstral Coefficients (MFCCs) derived from Short-Time Fourier Transform (STFT), offering a fixed time-frequency resolution.
- MFCCs were optimal for older Hidden Markov Models/Gaussian Mixture Models (HMM/GMM) ASR systems due to their low dimensionality and decorrelation.
- Modern ASR systems utilize Deep Neural Networks (DNNs), which can handle more complex, high-dimensional, and correlated features, making traditional feature extraction methods potentially suboptimal.
Purpose of the Study:
- To investigate whether multi-resolution speech analysis, varying time-frequency resolutions, can enhance acoustic modeling in DNN-based ASR.
- To re-evaluate the impact of time-frequency resolution trade-offs in speech analysis for contemporary ASR architectures.
Main Methods:
- Implemented multi-resolution speech representations by concatenating spectra computed with diverse time-frequency resolutions.
- Utilized the Kaldi toolkit and the TIMIT speech corpus for baseline and modified system experiments.
- Compared multi-resolution features against baseline single-resolution features and advanced post-processed, speaker-adapted features.
Main Results:
- Multi-resolution speech representations generally outperformed baseline single-resolution representations in ASR tasks.
- The hypothesis that multi-resolution analysis improves acoustic modeling with DNNs was supported by experimental findings.
- Combining multi-resolution features with highly processed and speaker-adapted features resulted in only marginal improvements over the best single-resolution methods.
Conclusions:
- Multi-resolution speech analysis offers a viable approach to potentially improve DNN-based ASR systems.
- The benefits of multi-resolution features may be less pronounced when integrated with highly optimized, advanced feature engineering techniques.
- Further research is warranted to fully leverage multi-resolution analysis within sophisticated ASR pipelines.
Related Concept Videos
Automatic Processing and Automatic Social Behavior
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Network Covalent Solids
To break or to melt a covalent network solid, covalent bonds must be broken. Because covalent bonds are relatively strong, covalent network solids are typically...
What is an Experiment?
Neural Regulation
Thomson's e/m Experiment
A particle with charge q, speed v, and mass m enters an area from the top, where the magnetic and electric fields are perpendicular both to the particle's motion and to one another. The magnetic...

