Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Perceiving Loudness, Pitch, and Location01:21

Perceiving Loudness, Pitch, and Location

218
The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
218
Unsoundness of Aggregate due to Volume Change01:26

Unsoundness of Aggregate due to Volume Change

115
Unsoundness in aggregates due to volume changes is primarily caused by the physical alterations aggregates undergo, such as freezing and thawing, thermal changes, and wetting and drying. Unsound aggregates, when subjected to these changes, result in volume change upon disintegration. This, in turn, contributes to the deterioration of concrete, including scaling, pop-outs, and cracking. Particular types of aggregates, such as porous flints, cherts, and those containing clay minerals, are...
115
Downsampling01:20

Downsampling

164
When considering a sampled sequence with zero values between sampling instants, one can replace it by taking every N-th value of the sequence. At these integer multiples of N, the original and sampled sequences coincide. This process, known as decimation, involves extracting every N-th sample from a sequence, thereby creating a more efficient sequence.
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...
164
Improving Translational Accuracy02:07

Improving Translational Accuracy

2.6K
2.6K
Sound Intensity Level00:53

Sound Intensity Level

4.2K
Humans perceive sound by hearing. The human ear helps sound waves reach the brain, which then interprets the waves and creates the perception of hearing. The loudness of the environment in which a person is located determines whether they can distinguish between different sound sources.
The human ear can perceive an extensive range of sound intensity, necessitating the use of the logarithmic scale to define a physical quantity—the intensity level. It is a ratio of two intensities and...
4.2K
Upsampling01:22

Upsampling

240
Managing signal sampling rates is essential in digital signal processing to maintain signal integrity. A decimated signal, characterized by a reduced frequency range due to its lower sampling rate, can be upsampled by inserting zeros between each sample. This upsampling process expands the original spectrum and introduces repeated spectral replicas at intervals dictated by the new Nyquist frequency. To refine this zero-inserted sequence, it is passed through a lowpass filter with a cutoff...
240

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

As group size increases, individuals modify their vocal features to signal cooperation while remaining recognizable.

Frontiers in psychology·2026
Same author

A multimodal speech-production dataset with time-aligned articulography, EEG, audio, and vocal-tract anatomy.

Scientific data·2026
Same author

Vocal-tract length estimation from vowel formants benchmarked against acoustic pharyngometry.

The Journal of the Acoustical Society of America·2026
Same author

Impaired Self-Other Voice Discrimination in Patients with Auditory-Verbal Hallucinations and Nonclinical Hallucination Proneness.

Schizophrenia bulletin·2026
Same author

Implicit voice learning through discrimination outperforms explicit listen-and-memorize tasks.

Scientific reports·2026
Same author

Acoustic and Perceptual Differences of <i>Aegyo</i> Speaking Style Across Gender in Seoul Korean.

Language and speech·2025

Related Experiment Video

Updated: Jul 12, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.5K

Acoustic compression in Zoom audio does not compromise voice recognition performance.

Valeriia Perepelytsia1, Volker Dellwo2

  • 1Department of Computational Linguistics, University of Zurich, Andreasstrasse 15, 8050, Zurich, Switzerland. valeriia.perepelytsia@uzh.ch.

Scientific Reports
|November 1, 2023
PubMed
Summary

Video conferencing audio quality, like Zoom, is as effective as studio audio for human voice recognition, outperforming lower-quality telephone audio. Familiarization with Zoom audio may even slightly improve voice recognition performance.

More Related Videos

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

464
Systematic Hearing Performance Evaluation Process for Adolescents with Cochlear Implantation at Early Ages
06:04

Systematic Hearing Performance Evaluation Process for Adolescents with Cochlear Implantation at Early Ages

Published on: March 24, 2023

410

Related Experiment Videos

Last Updated: Jul 12, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.5K
Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

464
Systematic Hearing Performance Evaluation Process for Adolescents with Cochlear Implantation at Early Ages
06:04

Systematic Hearing Performance Evaluation Process for Adolescents with Cochlear Implantation at Early Ages

Published on: March 24, 2023

410

Area of Science:

  • Speech processing
  • Human-computer interaction
  • Acoustic signal analysis

Background:

  • Human voice recognition accuracy is typically reduced over telephone channels compared to studio recordings.
  • Lossy compression in video conferencing affects audio quality and potentially voice recognition performance.

Purpose of the Study:

  • To investigate how video conferencing audio quality impacts human voice recognition.
  • To compare voice recognition performance across telephone, Zoom, and studio audio conditions.

Main Methods:

  • An old-new voice recognition task was employed.
  • Participants were familiarized with female voices in studio, Zoom, or telephone quality audio.
  • Voice recognition tests were conducted under matched and mismatched audio familiarization and testing conditions.

Main Results:

  • Voice recognition performance in Zoom audio was not significantly different from studio audio.
  • Both Zoom and studio audio yielded significantly better recognition than telephone audio.
  • Familiarization with Zoom audio showed a trend towards improved recognition performance compared to studio audio.

Conclusions:

  • Zoom's audio processing provides relevant information for voice recognition, comparable to studio quality.
  • Telephone audio quality significantly hinders voice recognition performance.
  • Further research may explore specific speech coding mechanisms in Zoom that could enhance voice recognition.