Related Experiment Videos
Identification of prosodic attitudes by a temporal recurrent network
Jean-Marc Blanc1, Peter Ford Dominey
1Institut des Sciences Cognitives, UMR 5015 CNRS, University Claude Bernard Lyon 1, 67 Boulevard Pinel, 69675 Bron Cedex, France.
Brain Research. Cognitive Brain Research
|October 17, 2003
Summary
This study shows that the temporal patterns in fundamental frequency (F0) can signal prosodic attitudes like surprise. A neural network model achieved higher accuracy than humans in decoding these attitudes from speech.
Area of Science:
- Computational neuroscience
- Linguistics
- Speech processing
Background:
- Human speech utilizes fundamental frequency (F0) modulation to convey prosodic attitudes.
- Decoding these prosodic attitudes from speech is crucial for natural language understanding.
- Previous research has explored various acoustic cues, but the temporal dynamics of F0 remain a key area of investigation.
Purpose of the Study:
- To investigate the role of temporal F0 structure in discriminating prosodic attitudes.
- To develop and evaluate a temporal recurrent neural network (TRN) for prosodic attitude classification.
- To compare the TRN's performance against human subject performance in a similar discrimination task.
Main Methods:
- A temporal recurrent neural network (TRN), inspired by primate frontostriatal neurophysiology, was employed.
- The TRN processed population coding of continuous, time-varying fundamental frequency (F0) values from natural language sentences.
- The model was trained and tested on an experiment involving discrimination of six prosodic attitudes.
Main Results:
- The TRN model achieved 82.52% accuracy in discriminating six prosodic attitudes.
- Human subjects achieved 72.8% accuracy in the same discrimination task.
- The results indicate that F0 temporal structure contains significant information for prosodic attitude classification.
Conclusions:
- Fundamental frequency (F0) variations over time are highly informative for identifying prosodic attitudes in speech.
- The developed temporal recurrent neural network (TRN) demonstrates effective categorical sensitivity to these F0 dynamics.
- The TRN model shows potential as a computational tool for understanding prosodic perception and classification.