Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Perceiving Loudness, Pitch, and Location01:21

Perceiving Loudness, Pitch, and Location

1.3K
The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
1.3K
Residuals and Least-Squares Property01:11

Residuals and Least-Squares Property

9.8K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
9.8K
Linear Approximation in Frequency Domain01:26

Linear Approximation in Frequency Domain

434
Linear systems are characterized by two main properties: superposition and homogeneity. Superposition allows the response to multiple inputs to be the sum of the responses to each individual input. Homogeneity ensures that scaling an input by a scalar results in the response being scaled by the same scalar.
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
434
Expected Frequencies in Goodness-of-Fit Tests01:19

Expected Frequencies in Goodness-of-Fit Tests

8.9K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n)  to the number of categories (k).
8.9K
Estimation of the Physical Quantities01:05

Estimation of the Physical Quantities

8.7K
On many occasions, physicists, other scientists, and engineers need to make estimates of a particular quantity. These are sometimes referred to as guesstimates, order-of-magnitude approximations, back-of-the-envelope calculations, or Fermi calculations. The physicist Enrico Fermi was famous for his ability to estimate various kinds of data with surprising precision. Estimating does not mean guessing a number or a formula at random. Instead, estimation means using prior experience and sound...
8.7K
Estimating Population Standard Deviation01:26

Estimating Population Standard Deviation

3.5K
When the population standard deviation is unknown and the sample size is large, the sample standard deviation s is commonly used as a point estimate of σ. However, it can sometimes under or overestimate the population standard deviation. To overcome this drawback, confidence intervals are determined to estimate population parameters and eliminate any calculation bias accurately. However, this only applies to random samples from normally distributed populations. Knowing the sample mean and...
3.5K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Deep learning-based environmental source separation and sound enhancement: Advancements for cochlear implant and normal hearing listeners.

The Journal of the Acoustical Society of America·2026
Same author

Capabilities of the CCi-MOBILE cochlear implant research platform for real-time sound coding.

The Journal of the Acoustical Society of America·2025
Same author

Speech Enhancement for Cochlear Implant Recipients using Deep Complex Convolution Transformer with Frequency Transformation.

IEEE/ACM transactions on audio, speech, and language processing·2025
Same author

Multi-objective non-intrusive hearing-aid speech assessment model.

The Journal of the Acoustical Society of America·2024
Same author

Advanced accent/dialect identification and accentedness assessment with multi-embedding models and automatic speech recognition.

The Journal of the Acoustical Society of America·2024
Same author

Child-adult speech diarization in naturalistic conditions of preschool classrooms using room-independent ResNet model and automatic speech recognition-based re-segmentation.

The Journal of the Acoustical Society of America·2024

Related Experiment Video

Updated: Apr 4, 2026

Asthma Detection Research Based on Voice Signal Processing and Machine Learning
04:04

Asthma Detection Research Based on Voice Signal Processing and Machine Learning

Published on: July 22, 2025

1.2K

Speaker height estimation from speech: Fusing spectral regression and statistical acoustic models.

John H L Hansen1, Keri Williams1, Hynek Bořil1

  • 1Center for Robust Speech Systems, Erik Jonsson School of Engineering and Computer Science, University of Texas at Dallas, Richardson, Texas 75083, USA.

The Journal of the Acoustical Society of America
|September 3, 2015
PubMed
Summary

This study introduces new methods for estimating speaker height using acoustic models and formant analysis. The research significantly reduces height estimation errors for both male and female speakers.

More Related Videos

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

980
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

2.1K

Related Experiment Videos

Last Updated: Apr 4, 2026

Asthma Detection Research Based on Voice Signal Processing and Machine Learning
04:04

Asthma Detection Research Based on Voice Signal Processing and Machine Learning

Published on: July 22, 2025

1.2K
Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

980
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

2.1K

Area of Science:

  • Speech processing
  • Forensic acoustics
  • Biometric analysis

Background:

  • Speaker height estimation is valuable for voice forensics and speaker recognition.
  • Previous methods have limitations in accuracy and real-world applicability.

Purpose of the Study:

  • To develop and evaluate novel statistical and formant analysis approaches for speaker height estimation.
  • To assess system performance, error sources, and robustness under various conditions.

Main Methods:

  • A statistical approach using Gaussian mixture models with acoustic models.
  • A formant analysis approach employing linear regression on selected speech sounds (phones).
  • Fusion of both systems and open-set testing on diverse audio datasets (TIMIT, MAR, YouTube).

Main Results:

  • The proposed algorithms demonstrate competitive performance against existing literature.
  • Mean average errors of 4.89 cm (males) and 4.55 cm (females) were achieved on the TIMIT corpus.
  • Open-set testing revealed the impact of channel mismatch on system performance.

Conclusions:

  • The developed algorithms offer a significant reduction in speaker height estimation error.
  • The study highlights the importance of considering channel variability in real-world applications.
  • These advancements contribute to more accurate voice forensic analysis and speaker identification.