Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Reducing Line Loss01:18

Reducing Line Loss

151
In a three-phase circuit, line loss is an indicator of energy dissipated as heat due to the resistance of transmission lines. To address this, incorporating transformers into the system—a step-up transformer at the source and a step-down transformer at the load—is a strategic solution. Two three-phase transformers are introduced to improve this.
With a step-up transformer at the source, the voltage is increased, thereby reducing the current in the transmission lines since power loss...
151
Difference from Background: Limit of Detection01:05

Difference from Background: Limit of Detection

6.4K
The limit of detection (LOD) is the smallest amount of analyte that can be distinguished from the background noise. The LOD value corresponds to the concentration at which the analyte signal is three times larger than the standard deviation of the blank signal. Below this value, the analyte signal cannot be differentiated from the background noise. It is calculated by dividing the calibration slope by 3 times the standard deviation of the blank signals.
The LOD indicates the presence or absence...
6.4K
Wilcoxon Signed-Ranks Test for Matched Pairs01:09

Wilcoxon Signed-Ranks Test for Matched Pairs

121
The Wilcoxon signed-rank test for matched pairs evaluates the null hypothesis by combining the ranks of differences with their signs. It essentially tests whether the median of the differences in a population of matched pairs is zero. Since the test incorporates more information than the sign test, it generally yields more trustable conclusions. This test also does not require the data to follow a normal distribution, but two conditions must be met for it to be applicable: (1) the data must...
121
Line Loss01:10

Line Loss

245
The different configurations of source-load connections include wye (star) and delta connections. The relationship between line and phase voltages and currents varies depending on the configuration. When the source is supplying power, it is transmitted through the wires to the load, and during this transmission, some power is absorbed by the wires, leading to line loss.
Line loss impacts power delivery efficiency in a balanced three-phase circuit. The symmetry in such a circuit simplifies the...
245
Residuals and Least-Squares Property01:11

Residuals and Least-Squares Property

7.4K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.4K
Expected Frequencies in Goodness-of-Fit Tests01:19

Expected Frequencies in Goodness-of-Fit Tests

2.5K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n)  to the number of categories (k).
2.5K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Novelty Detection in Underwater Acoustic Environments for Maritime Surveillance Using an Out-of-Distribution Detector for Neural Networks.

Sensors (Basel, Switzerland)·2026
Same author

Spatiotemporal Anomaly Detection in Distributed Acoustic Sensing Using a GraphDiffusion Model.

Sensors (Basel, Switzerland)·2025
Same author

Informer-Based Temperature Prediction Using Observed and Numerical Weather Prediction Data.

Sensors (Basel, Switzerland)·2023
Same author

End-to-End Model-Based Detection of Infants with Autism Spectrum Disorder Using a Pretrained Model.

Sensors (Basel, Switzerland)·2023
Same author

Two-Step Joint Optimization with Auxiliary Loss Function for Noise-Robust Speech Recognition.

Sensors (Basel, Switzerland)·2022
Same author

An Efficient Compression Method of Underwater Acoustic Sensor Signals for Underwater Surveillance.

Sensors (Basel, Switzerland)·2022

Related Experiment Video

Updated: Jun 27, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.5K

Cluster-Based Pairwise Contrastive Loss for Noise-Robust Speech Recognition.

Geon Woo Lee1, Hong Kook Kim1,2,3

  • 1AI Graduate School, Gwangju Institute of Science and Technology, Gwangju 61005, Republic of Korea.

Sensors (Basel, Switzerland)
|April 27, 2024
PubMed
Summary

This study introduces a joint training method for speech enhancement (SE) and automatic speech recognition (ASR) using a novel cluster-based pairwise contrastive (CBPC) loss. This approach improves speech quality and reduces word error rates (WER) in noisy conditions.

Keywords:
acoustic tokenizercontrastive lossjoint trainingnoise-robust speech recognitionself-supervised learningspeech enhancement

More Related Videos

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

443
Systematic Hearing Performance Evaluation Process for Adolescents with Cochlear Implantation at Early Ages
06:04

Systematic Hearing Performance Evaluation Process for Adolescents with Cochlear Implantation at Early Ages

Published on: March 24, 2023

380

Related Experiment Videos

Last Updated: Jun 27, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.5K
Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

443
Systematic Hearing Performance Evaluation Process for Adolescents with Cochlear Implantation at Early Ages
06:04

Systematic Hearing Performance Evaluation Process for Adolescents with Cochlear Implantation at Early Ages

Published on: March 24, 2023

380

Area of Science:

  • Artificial Intelligence
  • Speech Processing
  • Machine Learning

Background:

  • Joint training of speech enhancement (SE) and automatic speech recognition (ASR) models is crucial for improving speech processing performance.
  • Leveraging linguistic information from ASR to SE models can enhance SE capabilities.
  • Existing methods may not fully capture linguistic nuances or can overfit to noisy data.

Purpose of the Study:

  • To propose a novel joint training approach for SE and ASR models.
  • To introduce an acoustic tokenizer and a cluster-based pairwise contrastive (CBPC) loss function for effective linguistic information transfer.
  • To improve speech enhancement quality and reduce word error rate (WER) in noisy environments.

Main Methods:

  • A pipeline integrating SE and ASR models with an acoustic tokenizer was developed.
  • The acoustic tokenizer generated pseudo-labels using K-means clustering from ASR encoder outputs.
  • A novel cluster-based pairwise contrastive (CBPC) loss function, combined with infoNCE, was proposed for self-supervised learning to transfer linguistic information.

Main Results:

  • The proposed joint training approach with CBPC loss achieved a lower word error rate (WER) compared to conventional methods.
  • Speech quality scores were significantly higher than standalone SE models and those trained with traditional joint approaches.
  • The combination of CBPC and infoNCE loss demonstrated effectiveness in reducing WER and enhancing speech quality.

Conclusions:

  • The proposed CBPC loss function effectively transfers linguistic information from ASR to SE models within a joint training framework.
  • This approach leads to superior speech enhancement performance and more accurate speech recognition in noisy conditions.
  • The method offers a promising direction for advancing robust speech processing systems.