Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Perceiving Loudness, Pitch, and Location01:21

Perceiving Loudness, Pitch, and Location

160
The human brain perceives pitch through two primary mechanisms reflected in place theory and frequency theory. Each mechanism describes how sound waves are interpreted as specific pitches by the brain, offering insights into the intricate processes of auditory perception.
Place theory, or place coding, suggests that different pitches are heard because various sound waves activate specific locations along the cochlea's basilar membrane. The brain determines the pitch of a sound by...
160
Improving Translational Accuracy02:07

Improving Translational Accuracy

8.5K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
8.5K
Detection of Gross Error: The Q Test01:00

Detection of Gross Error: The Q Test

4.4K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
4.4K
The Ideal Transformer01:26

The Ideal Transformer

321
In single-phase two-winding transformers, two windings are coiled around a magnetic core characterized by cross-sectional area A and magnetic permeability μ. A phasor current i1 enters the left winding while i2 exits the right winding, establishing the fundamental working of the transformer through electromagnetic principles.
Ampere's Law forms the basis of understanding the magnetic field within the transformer. It states that the integral of the magnetic field intensity's...
321
Types Of Transformers01:16

Types Of Transformers

923
Transformers can provide desired voltages to a circuit by modifying the number of turns in the secondary windings.
If the ratio of the number of turns in the secondary winding to that of the primary winding is greater than one, then the transformer is said to be a step-up transformer. In a step-up transformer, the voltage at the secondary winding is greater than the voltage applied at the primary winding.
However, if this ratio is less than one, the transformer is said to be a step-down...
923
Transformers with Off-Nominal Turns Ratios01:25

Transformers with Off-Nominal Turns Ratios

122
In scenarios involving parallel transformers with disparate ratings, developing per-unit models requires accommodating off-nominal turns ratios. This situation arises when the selected base voltages are not proportional to the transformer’s voltage ratings. Consider a transformer where the rated voltages are related by the term a. If the chosen voltage bases satisfy a relationship involving term b, term c is defined as the ratio of these bases. This ratio is then substituted into the...
122

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Preterm labour induction: modalities, implications and outcomes.

European journal of obstetrics, gynecology, and reproductive biology·2025
Same author

Exploring crystallized and fluid intelligence in down syndrome using graph theory.

Scientific reports·2024
Same author

Classifying interpersonal synchronization states using a data-driven approach: implications for social interaction understanding.

Scientific reports·2023
Same author

Intuitive Cognition-Based Method for Generating Speech Using Hand Gestures.

Sensors (Basel, Switzerland)·2021
Same author

Save Our Roads from GNSS Jamming: A Crowdsource Framework for Threat Evaluation.

Sensors (Basel, Switzerland)·2021
Same author

Impairments of interpersonal synchrony evident in attention deficit hyperactivity disorder (ADHD).

Acta psychologica·2020

Related Experiment Video

Updated: May 10, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.3K

Transformer-based language-independent gender recognition in noisy audio environments.

Or Haim Anidjar1,2, Roi Yozevitch3,4

  • 1Faculty of Computer Science, College of Management, Rishon Le'Tzion, Israel.

Scientific Reports
|April 25, 2025
PubMed
Summary

This study introduces a novel method for speaker gender identification in noisy audio. Mel-spectrogram analysis outperformed the Wav2Vec2 acoustic model, achieving 99% accuracy in Russian, and highlighting equitable dataset needs.

Keywords:
Automatic speech recognitionLanguage independent gender recognitionWav2Vec 2.0

More Related Videos

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

377
Using Eye Movements Recorded in the Visual World Paradigm to Explore the Online Processing of Spoken Language
09:27

Using Eye Movements Recorded in the Visual World Paradigm to Explore the Online Processing of Spoken Language

Published on: October 13, 2018

9.8K

Related Experiment Videos

Last Updated: May 10, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.3K
Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
09:09

Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody

Published on: September 27, 2024

377
Using Eye Movements Recorded in the Visual World Paradigm to Explore the Online Processing of Spoken Language
09:27

Using Eye Movements Recorded in the Visual World Paradigm to Explore the Online Processing of Spoken Language

Published on: October 13, 2018

9.8K

Area of Science:

  • Speech processing
  • Machine learning
  • Computational linguistics

Background:

  • Speaker gender identification is crucial for voice recognition systems.
  • Existing systems often exhibit gender bias due to imbalanced training data.
  • Noise and language variations pose significant challenges in accurate gender detection.

Purpose of the Study:

  • To develop an independent method for identifying speaker gender from audio clips in noisy environments.
  • To compare the effectiveness of Mel-spectrograms versus Wav2Vec2 acoustic model emissions for gender identification.
  • To address and mitigate gender bias in voice recognition by ensuring balanced datasets.

Main Methods:

  • Audio clips were processed using Mel-spectrograms and Wav2Vec2 acoustic model emissions.
  • Experiments were conducted across five languages: English, Arabic, Spanish, French, and Russian.
  • A balanced dataset with equivalent male and female audio clips was used to mitigate bias.

Main Results:

  • The Mel-spectrogram method demonstrated superior performance over the Wav2Vec2 transformer method.
  • Spectrogram analysis achieved 99% accuracy for Russian, while Wav2Vec2 achieved 89%.
  • Models trained on diverse languages and both noisy/silent conditions showed improved accuracy.

Conclusions:

  • Mel-spectrograms offer a robust approach for acoustic gender detection, outperforming transformer-based models in this study.
  • Balanced datasets are essential for reducing gender bias in speaker recognition systems.
  • Future systems should incorporate multi-lingual and varied environmental data for enhanced reliability and fairness.