Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Sample Handling01:02

Sample Handling

2.8K
Transportation of samples from the collection point to the laboratory, as well as storage and preservation techniques, are crucial for maintaining sample integrity and ensuring accurate and reliable test results.
Samples should be transported carefully from collection points to the laboratory. They should be properly sealed and clearly labeled to prevent cross-contamination. To preserve the sample integrity, optimal temperature conditions during transport are essential. This could involve using...
2.8K
Sound Intensity Level00:53

Sound Intensity Level

5.0K
Humans perceive sound by hearing. The human ear helps sound waves reach the brain, which then interprets the waves and creates the perception of hearing. The loudness of the environment in which a person is located determines whether they can distinguish between different sound sources.
The human ear can perceive an extensive range of sound intensity, necessitating the use of the logarithmic scale to define a physical quantity—the intensity level. It is a ratio of two intensities and...
5.0K
Sign Test for Nominal Data01:12

Sign Test for Nominal Data

429
The sign test is a nonparametric method used to evaluate hypotheses about the median of a single sample or to compare the medians of two related samples. The sign test is particularly useful when dealing with nominal data, which includes distinct categories without an inherent order, such as names, labels, and preferences. Nominal data restricts statistical analysis to evaluating population proportions rather than mean or median values that require continuous data.
For example, consider a...
429
Sound Intensity00:58

Sound Intensity

5.0K
The loudness of a sound source is related to how energetically the source is vibrating, consequently making the molecules of the propagation medium vibrate. To measure the loudness of a source, the physical quantity of interest is the intensity. This is defined as the energy emitted per unit of time per unit of area perpendicular to the sound wave's propagation direction. Since the total energy is greater if the source vibrates for a longer duration and over a larger area, dividing the...
5.0K
Perceiving Loudness, Pitch, and Location01:21

Perceiving Loudness, Pitch, and Location

1.3K
The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
1.3K
Hearing01:31

Hearing

58.4K
When we hear a sound, our nervous system is detecting sound waves—pressure waves of mechanical energy traveling through a medium. The frequency of the wave is perceived as pitch, while the amplitude is perceived as loudness.
58.4K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Under-resourced dialect identification in Ao using source information.

The Journal of the Acoustical Society of America·2022
Same author

Overlapped speech detection using phase features.

The Journal of the Acoustical Society of America·2021
Same author

Analysis and modeling of dialect information in Ao, a low resource language.

The Journal of the Acoustical Society of America·2021
Same author

Exploration of excitation source information for shouted and normal speech classification.

The Journal of the Acoustical Society of America·2020
Same author

Detection and assessment of hypernasality in repaired cleft palate speech using vocal tract and residual features.

The Journal of the Acoustical Society of America·2020
Same author

Objective assessment of cleft lip and palate speech intelligibility using articulation and hypernasality measures.

The Journal of the Acoustical Society of America·2019

Related Experiment Video

Updated: Mar 17, 2026

Sound Source Localization Testing in Single-sided Deafness Following Bone Conduction Intervention
04:32

Sound Source Localization Testing in Single-sided Deafness Following Bone Conduction Intervention

Published on: December 20, 2024

958

Exploring different attributes of source information for speaker verification with limited test data.

Rohan Kumar Das1, S R Mahadeva Prasanna1

  • 1Department of Electronics and Electrical Engineering, Indian Institute of Technology Guwahati, Guwahati-781039, Assam, India.

The Journal of the Acoustical Society of America
|August 1, 2016
PubMed
Summary

This study enhances speaker verification using novel features like mel power difference and cepstral coefficients. Combining these improved accuracy, especially with limited test data.

More Related Videos

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

2.1K
Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

957

Related Experiment Videos

Last Updated: Mar 17, 2026

Sound Source Localization Testing in Single-sided Deafness Following Bone Conduction Intervention
04:32

Sound Source Localization Testing in Single-sided Deafness Following Bone Conduction Intervention

Published on: December 20, 2024

958
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

2.1K
Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

957

Area of Science:

  • Speech processing
  • Biometrics
  • Signal analysis

Background:

  • Speaker verification systems often struggle with limited test data.
  • Traditional features like Mel Frequency Cepstral Coefficients (MFCCs) may not capture all relevant speaker-specific information.

Purpose of the Study:

  • To investigate the effectiveness of novel acoustic features for speaker verification under data scarcity.
  • To combine multiple source features to improve system performance.

Main Methods:

  • Exploration of mel power difference of spectrum in subband.
  • Analysis of residual mel frequency cepstral coefficient.
  • Application of Discrete Cosine Transform (DCT) on integrated linear prediction residual.
  • Feature combination and evaluation on the NIST SRE 2003 database.

Main Results:

  • The proposed combination of three source features (mel power difference, residual MFCC, DCT of linear prediction residual) achieved a lower Equal Error Rate (EER) of 20.19% and Decision Cost Function (DCF) of 0.3759.
  • This performance surpasses the baseline MFCC feature, which yielded an EER of 22.31% and DCF of 0.4128.
  • The individual features capture distinct speaker attributes: periodicity, smoothed spectrum, and glottal signal shape.

Conclusions:

  • Combining mel power difference, residual MFCC, and DCT of linear prediction residual offers a robust approach to speaker verification, particularly effective under limited test data conditions.
  • These complementary features provide a more comprehensive representation of speaker-specific characteristics than traditional MFCCs alone.