Related Experiment Video
Updated: Jan 14, 2026

05:48
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
2.0K
Accurate semi-supervised automatic speech recognition for ordinary and characterized speeches via
Ka Hyun Park1, Junghun Kim2, U Kang1
1Department of CSE, Seoul National University, Seoul, Republic of Korea.
Plos One
|October 21, 2025
Summary
This study introduces MOCA and MOCA-S, novel semi-supervised methods for Automatic Speech Recognition (ASR). These models enhance transcription accuracy for both ordinary and characterized speech by reducing reliance on potentially inaccurate pseudo-labels.
Area of Science:
- Artificial Intelligence
- Speech Processing
- Machine Learning
Background:
- Automatic Speech Recognition (ASR) systems are crucial for applications like translation and transcription.
- Current ASR models are specialized for either ordinary or characterized speech.
- Semi-supervised learning is gaining traction due to the high cost and scarcity of labeled speech data.
Purpose of the Study:
- To develop accurate semi-supervised ASR models for both ordinary and characterized speech.
- To address the limitations of pseudo-labeling in previous semi-supervised ASR approaches.
- To improve ASR performance, especially for characterized speech with limited data.
Main Methods:
- Proposed MOCA (Multi-hypotheses-based Curriculum learning for semi-supervised Asr) for ordinary speech.
- Developed MOCA-S for characterized speech, leveraging data from other speech types.
- MOCA and MOCA-S generate multiple hypotheses per instance to mitigate pseudo-label dependency.
- MOCA-S dynamically adjusts pseudo-labeling based on trait relevance.
Main Results:
- MOCA and MOCA-S demonstrated significant improvements in ASR accuracy compared to existing models.
- The multi-hypotheses approach effectively reduced reliance on inaccurate pseudo labels.
- MOCA-S successfully utilized data from other speech traits to enhance characterized speech recognition.
Conclusions:
- The proposed MOCA and MOCA-S frameworks offer a robust solution for semi-supervised ASR.
- These methods enhance transcription accuracy for diverse speech types, including those with unique characteristics.
- The study highlights the effectiveness of multi-hypotheses generation and cross-trait data utilization in ASR.
Related Concept Videos
Automatic Processing and Automatic Social Behavior
213
Automatic processing refers to the cognitive operations that occur without conscious intent or awareness, playing a fundamental role in shaping social cognition and behavior. These processes enable individuals to navigate complex social environments efficiently by relying on mental shortcuts and pre-existing knowledge structures known as schemas. One of the most influential mechanisms underlying automatic processing is priming, which subtly activates mental representations through exposure to...
213
Multi-input and Multi-variable systems
385
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence of...
In the absence of...
385
Classification of Signals
1.3K
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
1.3K

