Related Experiment Video
Updated: Jun 11, 2025

05:48
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
1.4K
FluencyBank Timestamped: An Updated Data Set for Disfluency Detection and Automatic Intended Speech Recognition
Amrit Romana1, Minxue Niu1, Matthew Perez1
1University of Michigan.
Journal of Speech, Language, and Hearing Research : JSLHR
|October 8, 2024
Summary
This study updates the FluencyBank dataset with timestamps and disfluency labels, revealing performance gaps in speech recognition and disfluency detection for people who stutter (PWS). The new FluencyBank Timestamped dataset aims to improve speech processing models for PWS.
Area of Science:
- Computational Linguistics
- Speech Processing
- Human-Computer Interaction
Background:
- Speech processing models often perform less accurately on speech from people who stutter (PWS).
- Existing datasets may not adequately represent the nuances of disfluent speech.
- There is a need for specialized datasets to benchmark and improve models for diverse speech patterns.
Purpose of the Study:
- To introduce FluencyBank Timestamped, an updated dataset with precise word timings and disfluency annotations.
- To enable comparative analysis of speech processing models on typical versus stuttered speech.
- To benchmark the performance of speech recognition and disfluency detection models on PWS speech.
Main Methods:
- Updated FluencyBank dataset with semi-automated, manually reviewed transcripts, timestamps, and disfluency labels.
- Evaluated Whisper for intended speech recognition and BERT/Whisper for disfluency detection.
- Compared model performance on Switchboard (typical speech) and FluencyBank Timestamped (PWS speech).
Main Results:
- Intended speech word error rate (isWER) was comparable between datasets, but Whisper transcribed filled pauses and partial words more frequently in FluencyBank Timestamped.
- isWER increased with stuttering severity within FluencyBank Timestamped.
- Models struggled with repetition detection in PWS speech, indicating generalization issues.
Conclusions:
- Significant performance gaps exist between speech processing models for typical speech and speech from PWS.
- FluencyBank Timestamped is a valuable resource for researchers aiming to close these performance gaps.
- Further advancements are needed to develop robust speech processing models that perform equitably across all speakers.

