Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same journal

RETRACTED: Ndaguba et al. Operability of Smart Spaces in Urban Environments: A Systematic Review on Enhancing Functionality and User Experience. <i>Sensors</i> 2023, <i>23</i>, 6938.

Sensors (Basel, Switzerland)·2026
Same journal

Correction: Ma et al. A Lightweight, Low-Frequency, Broadband Underwater Acoustic Transducer with Ternary Symmetric Excitation: Integrating KNN and Terfenol-D for Enhanced Performance. <i>2026</i>, <i>26</i>, 3645.

Sensors (Basel, Switzerland)·2026
Same journal

Correction: He et al. An Edge-Computing-Based Emotion-Aware Adaptive Lighting System for Intelligent Cockpits. <i>Sensors</i> 2026, <i>26</i>, 3489.

Sensors (Basel, Switzerland)·2026
Same journal

Correction: Tu et al. Lower Limb Motion Recognition with Improved SVM Based on Surface Electromyography. <i>Sensors</i> 2024, <i>24</i>, 3097.

Sensors (Basel, Switzerland)·2026
Same journal

Real-Time Detection System for Road Roughness Based on Ultrasonic Technology.

Sensors (Basel, Switzerland)·2026
Same journal

FedHSFV: Federated Learning for Finger Vein Recognition via Hierarchical Decoupling and Subspace Metric.

Sensors (Basel, Switzerland)·2026

Related Experiment Video

Updated: Aug 16, 2025

A Method to Study Adaptation to Left-Right Reversed Audition
07:14

A Method to Study Adaptation to Left-Right Reversed Audition

Published on: October 29, 2018

6.6K

Domain Adaptation with Augmented Data by Deep Neural Network Based Method Using Re-Recorded Speech for Automatic

Raufun Nahar1, Shogo Miwa2, Atsuhiko Kai1

  • 1Graduate School of Science and Technology, Shizuoka University, Hamamatsu 432-8561, Japan.

Sensors (Basel, Switzerland)
|December 23, 2022
PubMed
Summary

This study enhances automatic speech recognition (ASR) by using simulated data for training. The proposed method significantly reduces errors for real-world speech, especially over Long Term Evolution (LTE) channels.

Keywords:
ASRDNNVoLTEclassroom recordingdata augmentationfeature transformationreal environmentrecording alignment

More Related Videos

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.6K
Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

511

Related Experiment Videos

Last Updated: Aug 16, 2025

A Method to Study Adaptation to Left-Right Reversed Audition
07:14

A Method to Study Adaptation to Left-Right Reversed Audition

Published on: October 29, 2018

6.6K
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.6K
Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

511

Area of Science:

  • Speech Recognition
  • Artificial Intelligence
  • Machine Learning

Background:

  • Automatic Speech Recognition (ASR) models, primarily based on Artificial Neural Networks (ANNs), require extensive, condition-matched training data.
  • Real-world speech data acquisition across diverse acoustic environments (e.g., mobile telephony, classroom recordings) is costly and challenging.
  • Existing ASR systems struggle with performance degradation due to variations in recording channels and environments.

Purpose of the Study:

  • To investigate and propose effective methods for domain adaptation in ASR using simulated and augmented data.
  • To reduce the cost and effort associated with acquiring diverse, real-world speech data for ASR training.
  • To improve the character error rate (CERR) of ASR systems in challenging acoustic conditions, specifically targeting telephone speech via Long Term Evolution (LTE) channels.

Main Methods:

  • Training ASR models with simulated augmented data, followed by fine-tuning for domain adaptation using deep neural network (DNN)-based simulated data and re-recorded data.
  • Employing DNN-based feature transformation to generate realistic speech features from clean condition recordings.
  • Conducting a comparative investigation of different recording channel adaptation techniques for real-world speech recognition.

Main Results:

  • The proposed method achieved a 27.0% character error rate reduction (CERR) for DNN-hidden Markov model (DNN-HMM) hybrid ASR approaches.
  • A significant 36.4% CERR was observed for end-to-end ASR approaches targeting the LTE channel of telephone speech.
  • Simulated data augmentation and DNN-based feature transformation proved effective for domain adaptation in ASR.

Conclusions:

  • Training ASR models with simulated data and fine-tuning with DNN-based techniques offers a cost-effective solution for improving performance in diverse acoustic environments.
  • The proposed approach demonstrates substantial improvements in ASR accuracy for real-world telephone speech transmitted over LTE channels.
  • This research provides a viable strategy for enhancing ASR robustness against acoustic variations without extensive real-world data collection.