Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Perceiving Loudness, Pitch, and Location01:21

Perceiving Loudness, Pitch, and Location

1.2K
The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
1.2K
Classification of Signals01:30

Classification of Signals

1.5K
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
1.5K
Hearing01:31

Hearing

58.0K
When we hear a sound, our nervous system is detecting sound waves—pressure waves of mechanical energy traveling through a medium. The frequency of the wave is perceived as pitch, while the amplitude is perceived as loudness.
58.0K
Auditory Pathway01:15

Auditory Pathway

7.8K
Auditory pathways constitute the complex neural circuits responsible for transmitting and interpreting auditory information from the peripheral auditory system to the brain. Sound waves are initially captured by the outer ear, funneled through the ear canal, and reach the tympanic membrane (eardrum). These vibrations are transmitted via the middle ear's ossicles to the inner ear's cochlea.
When viewed cross-sectionally, the cochlea reveals the scala vestibuli and scala tympani flanking...
7.8K
The Cochlea01:13

The Cochlea

51.8K
The cochlea is a coiled structure in the inner ear that contains hair cells—the sensory receptors of the auditory system. Sound waves are transmitted to the cochlea by small bones attached to the eardrum called the ossicles, which vibrate the oval window that leads to the inner ear. This causes fluid in the chambers of the cochlea to move, vibrating the basilar membrane.
51.8K
Linear Approximation in Frequency Domain01:26

Linear Approximation in Frequency Domain

407
Linear systems are characterized by two main properties: superposition and homogeneity. Superposition allows the response to multiple inputs to be the sum of the responses to each individual input. Homogeneity ensures that scaling an input by a scalar results in the response being scaled by the same scalar.
In contrast, nonlinear systems do not inherently possess these properties. However, for small deviations around an operating point, a nonlinear system can often be approximated as linear....
407

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Residual Neural Network precisely quantifies dysarthria severity-level based on short-duration speech segments.

Neural networks : the official journal of the International Neural Network Society·2021
Same author

Data Collection of Infant Cries for Research and Analysis.

Journal of voice : official journal of the Voice Foundation·2016
See all related articles

Related Experiment Video

Updated: Feb 28, 2026

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

2.0K

Auditory feature representation using convolutional restricted Boltzmann machine and Teager energy operator for

Hardik B Sailor1, Hemant A Patil1

  • 1Speech Research Lab, Dhirubhai Ambani Institute of Information and Communication Technology (DA-IICT), Gandhinagar-382007, Gujarat, India sailor_hardik@daiict.ac.in, hemant_patil@daiict.ac.in.

The Journal of the Acoustical Society of America
|June 17, 2017
PubMed
Summary

This study introduces novel auditory features using a learned filterbank and Teager energy operator (TEO) for improved speech recognition. These features outperform traditional Mel filterbanks, significantly reducing word error rates in noisy conditions.

More Related Videos

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

932
Machine Learning-Based Cough Tone Classification: Diagnostic Exploration of Chronic Obstructive Pulmonary Disease and Respiratory Tract Infections
06:22

Machine Learning-Based Cough Tone Classification: Diagnostic Exploration of Chronic Obstructive Pulmonary Disease and Respiratory Tract Infections

Published on: September 19, 2025

581

Related Experiment Videos

Last Updated: Feb 28, 2026

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

2.0K
Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

932
Machine Learning-Based Cough Tone Classification: Diagnostic Exploration of Chronic Obstructive Pulmonary Disease and Respiratory Tract Infections
06:22

Machine Learning-Based Cough Tone Classification: Diagnostic Exploration of Chronic Obstructive Pulmonary Disease and Respiratory Tract Infections

Published on: September 19, 2025

581

Area of Science:

  • Signal Processing
  • Machine Learning
  • Speech Recognition

Background:

  • Traditional auditory feature representations often struggle with noisy environments.
  • Mel filterbank features are widely used but have limitations in robustness.

Purpose of the Study:

  • To propose a new auditory feature representation technique for enhanced speech recognition.
  • To improve noise robustness in speech feature extraction.

Main Methods:

  • A filterbank was learned using an annealing dropout convolutional restricted Boltzmann machine (ConvRBM).
  • Noise-robust energy estimation was performed using the Teager energy operator (TEO) on subbands.
  • Short-term spectral features were obtained by pooling TEO-processed subbands.

Main Results:

  • The proposed features demonstrated superior performance compared to Mel filterbank features on the AURORA 4 database.
  • Relative improvements in word error rate ranged from 2.59% to 11.63% with time delay neural networks.
  • Relative improvements in word error rate ranged from 1.26% to 6.87% with bidirectional long short-term memory models.

Conclusions:

  • The proposed auditory feature representation technique offers significant advantages over Mel filterbanks.
  • The combination of ConvRBM-learned filterbanks and TEO enhances noise robustness in speech recognition.
  • This method shows promise for improving the performance of speech recognition systems in challenging acoustic conditions.