Related Experiment Video
Updated: Jul 7, 2026

05:38
Interaction between Phonological and Semantic Processes in Visual Word Recognition using Electrophysiology
Published on: June 29, 2021
A novel objective function for improved phoneme recognition using time-delay neural networks
1Carnegie-Mellon Univ., Pittsburgh, PA.
IEEE Transactions on Neural Networks
|January 1, 1990
Summary
A new training method for time-delay neural networks (TDNNs), the classification figure of merit (CFM), significantly reduces misclassifications in speech recognition. This approach improves accuracy for voice-stop consonants like /b,d,g/.
Area of Science:
- Speech Recognition
- Artificial Intelligence
- Machine Learning
Background:
- Traditional speech recognition models often use mean-squared-error (MSE) or cross-entropy (CE) objective functions for training.
- These functions aim to minimize discrepancies between predicted and actual outputs, which may not be optimal for classification tasks.
Purpose of the Study:
- To introduce and evaluate a novel objective function, the classification figure of merit (CFM), for training time-delay neural networks (TDNNs).
- To assess the effectiveness of CFM in improving the accuracy of single-speaker and multispeaker recognition of voice-stop consonants (/b,d,g/).
Main Methods:
- Implemented time-delay neural networks (TDNNs) for speech recognition tasks.
- Developed and applied a new objective function, the classification figure of merit (CFM), during TDNN training.
- Compared CFM against traditional MSE and cross-entropy (CE) objective functions, incorporating a simple arbitration mechanism.
Main Results:
- The CFM objective function demonstrated a median 30% reduction in misclassifications compared to TDNNs trained solely with the MSE objective function.
- CFM focuses on maximizing the difference for incorrect classifications, unlike MSE and CE which minimize overall error.
Conclusions:
- The classification figure of merit (CFM) offers a more effective approach for training TDNNs in speech recognition compared to traditional methods.
- This enhancement leads to significant improvements in the accurate recognition of challenging speech sounds like voice-stop consonants.
