Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Interference: Path Lengths01:10

Interference: Path Lengths

1.5K
Consider two sources of sound, that may or may not be in phase, emitting waves at a single frequency, and consider the frequencies to be the same.
Two special sources may be considered when they are in phase. This can be easily achieved by feeding the two sources from the same source. An example would be synchronizing the two speakers by feeding them with the same source, such as the sound waves produced by a tuning fork. This setup ensures that the two sources have the same frequency and are...
1.5K
¹H NMR: Interpreting Distorted and Overlapping Signals01:02

¹H NMR: Interpreting Distorted and Overlapping Signals

1.2K
Spin systems where the difference in chemical shifts of the coupled nuclei is greater than ten times J are called first-order spin systems. These nuclei are weakly coupled, and their chemical shifts and coupling constant can generally be estimated from the well-separated signals in the spectrum.
As Δν decreases and the signals move closer, the doublets appear increasingly distorted. The intensities of the inner lines increase at the cost of those of the outer lines as the signals are...
1.2K
Time and frequency -Domain Interpretation of Phase-lag Control01:21

Time and frequency -Domain Interpretation of Phase-lag Control

161
Phase-lag controllers are widely used in control systems to improve stability and reduce steady-state errors. A dimmer switch controlling the brightness of a light bulb serves as a practical example of phase-lag control, gradually adjusting the bulb's brightness. Mathematically, phase-lag control or low-pass filtering is represented when the factor 'a' is less than 1.
Phase-lag controllers do not place a pole at zero, but instead influence the steady-state error by amplifying any...
161
Time and frequency -Domain Interpretation of Phase-lead Control01:24

Time and frequency -Domain Interpretation of Phase-lead Control

154
Phase-lead controllers are commonly used in various control systems to enhance response speed and stability. Adjusting the brightness on a television screen offers a practical example of phase-lead control. When contrast is enhanced, a phase-lead controller is employed. Mathematically, phase-lead control is identified when the first parameter is smaller than the second.
The design of phase-lead control involves the strategic placement of poles and zeros to balance steady-state error and system...
154
Difference from Background: Limit of Detection01:05

Difference from Background: Limit of Detection

7.4K
The limit of detection (LOD) is the smallest amount of analyte that can be distinguished from the background noise. The LOD value corresponds to the concentration at which the analyte signal is three times larger than the standard deviation of the blank signal. Below this value, the analyte signal cannot be differentiated from the background noise. It is calculated by dividing the calibration slope by 3 times the standard deviation of the blank signals.
The LOD indicates the presence or absence...
7.4K
Frequency-Domain Interpretation of PD Control01:24

Frequency-Domain Interpretation of PD Control

187
Proportional-Derivative (PD) controllers are widely used in fan control systems to improve stability and performance. A fan control system can be effectively represented using a Bode plot to illustrate the impact of a PD controller through its transfer function. The Bode plot visually conveys how PD control modifies the fan's response across various frequencies, providing a frequency domain interpretation of the controller's behavior.
The proportional control gain, combined with the...
187

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Classification of microplanktons in an imbalanced digital holographic image dataset with a deep network using channel attention.

Journal of the Optical Society of America. A, Optics, image science, and vision·2025
Same author

Under-resourced dialect identification in Ao using source information.

The Journal of the Acoustical Society of America·2022
Same author

Analysis and modeling of dialect information in Ao, a low resource language.

The Journal of the Acoustical Society of America·2021
Same author

Exploration of excitation source information for shouted and normal speech classification.

The Journal of the Acoustical Society of America·2020
Same author

Detection and assessment of hypernasality in repaired cleft palate speech using vocal tract and residual features.

The Journal of the Acoustical Society of America·2020
Same author

Objective assessment of cleft lip and palate speech intelligibility using articulation and hypernasality measures.

The Journal of the Acoustical Society of America·2019

Related Experiment Video

Updated: Oct 15, 2025

Author Spotlight: Unlocking New Insights in fNIRS Studies - A Novel Framework for Inter-Brain Synchrony Analysis
05:59

Author Spotlight: Unlocking New Insights in fNIRS Studies - A Novel Framework for Inter-Brain Synchrony Analysis

Published on: October 6, 2023

2.8K

Overlapped speech detection using phase features.

Shikha Baghel1, S R Mahadeva Prasanna2, Prithwijit Guha1

  • 1Department of Electronics and Electrical Engineering, Indian Institute of Technology Guwahati, Guwahati-781039, India.

The Journal of the Acoustical Society of America
|October 31, 2021
PubMed
Summary

This study introduces novel phase-based features, Instantaneous Frequency Cosine Coefficient (IFCC) and Modified Group Delay Cepstral Coefficient (MGDCC), for overlapped speech detection. Combining these phase features with magnitude features significantly improves accuracy in recognizing simultaneous speech.

More Related Videos

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

584
Interaction between Phonological and Semantic Processes in Visual Word Recognition using Electrophysiology
05:38

Interaction between Phonological and Semantic Processes in Visual Word Recognition using Electrophysiology

Published on: June 29, 2021

2.5K

Related Experiment Videos

Last Updated: Oct 15, 2025

Author Spotlight: Unlocking New Insights in fNIRS Studies - A Novel Framework for Inter-Brain Synchrony Analysis
05:59

Author Spotlight: Unlocking New Insights in fNIRS Studies - A Novel Framework for Inter-Brain Synchrony Analysis

Published on: October 6, 2023

2.8K
Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

584
Interaction between Phonological and Semantic Processes in Visual Word Recognition using Electrophysiology
05:38

Interaction between Phonological and Semantic Processes in Visual Word Recognition using Electrophysiology

Published on: June 29, 2021

2.5K

Area of Science:

  • Signal Processing
  • Machine Learning
  • Acoustic Event Detection

Background:

  • Overlapped speech, where multiple speakers talk simultaneously, poses significant challenges for automatic speech recognition and speaker diarization systems.
  • Traditional methods often rely on magnitude spectrum features, potentially overlooking crucial phase information within the audio signal.

Purpose of the Study:

  • To investigate the efficacy of underutilized signal phase information for overlapped speech detection.
  • To develop and evaluate a novel system leveraging phase-based features for improved recognition of simultaneous speech.

Main Methods:

  • Exploration of Instantaneous Frequency Cosine Coefficient (IFCC) and Modified Group Delay Cepstral Coefficient (MGDCC) as novel phase-based features.
  • Implementation of a Convolutional Neural Network and Long Short-Term Memory (CNN-LSTM) classifier for overlapped speech detection.
  • Benchmarking against baseline methods using magnitude spectrum features on synthetically generated (GRID corpus) and real-world (AMI corpus) overlapped speech data.

Main Results:

  • The combination of IFCC and MGDCC features with the CNN-LSTM classifier demonstrated superior performance compared to baseline approaches.
  • Integrating phase features (IFCC, MGDCC) with magnitude-based Mel-frequency cepstral coefficients (MFCC) yielded the best overall performance, highlighting the value of complementary information.
  • The study analyzed the impact of segment duration, speaker gender, and the number of simultaneous speakers on detection accuracy.

Conclusions:

  • Signal phase information, particularly when combined with magnitude features, is crucial for enhancing overlapped speech detection systems.
  • The proposed CNN-LSTM model utilizing IFCC and MGDCC features offers a promising advancement for accurately identifying simultaneous speech.
  • The findings underscore the importance of exploring diverse feature sets to overcome limitations in current speech processing technologies.