Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Multi-input and Multi-variable systems01:22

Multi-input and Multi-variable systems

203
Cruise control systems in cars are designed as multi-input systems to maintain a driver's desired speed while compensating for external disturbances such as changes in terrain. The block diagram for a cruise control system typically includes two main inputs: the desired speed set by the driver and any external disturbances, such as the incline of the road. By adjusting the engine throttle, the system maintains the vehicle's speed as close to the desired value as possible.
In the absence...
203
Classification of Signals01:30

Classification of Signals

1.0K
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
1.0K
Force Classification01:22

Force Classification

1.8K
Forces play a crucial role in the study of physics and engineering. They are essential in describing the motion, behavior, and equilibrium of objects in the physical world. Forces can be classified based on their origin, type, and direction of action.
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
1.8K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

A Novel Method to Inspect 3D Ball Joint Socket Products Using 2D Convolutional Neural Network with Spatial and Channel Attention.

Sensors (Basel, Switzerland)·2022
Same author

COVID-19: Were Public Health Interventions and the Disclosure of Patients' Contact History Effective in Upholding Social Distancing? Evidence from South Korea.

Journal of multidisciplinary healthcare·2021
Same author

A CNN-Assisted Enhanced Audio Signal Processing for Speech Emotion Recognition.

Sensors (Basel, Switzerland)·2020
See all related articles

Related Experiment Video

Updated: Oct 20, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.7K

Age and Gender Recognition Using a Convolutional Neural Network with a Specially Designed Multi-Attention Module

Anvarjon Tursunov1, Mustaqeem1, Joon Yeon Choeh2

  • 1Interaction Technology Laboratory, Department of Software, Sejong University, Seoul 05006, Korea.

Sensors (Basel, Switzerland)
|September 10, 2021
PubMed
Summary

This study introduces a novel convolutional neural network (CNN) with a multi-attention module (MAM) for accurate speaker age and gender recognition from speech signals. The model achieves high accuracy in classifying gender, age, and age-gender from diverse datasets.

Keywords:
age and gender recognitionconvolutional neural networkhuman-computer interactionmulti-attention modulespeech signals

More Related Videos

Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention
06:37

Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention

Published on: December 15, 2023

4.5K

Related Experiment Videos

Last Updated: Oct 20, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.7K
Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention
06:37

Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention

Published on: December 15, 2023

4.5K

Area of Science:

  • Speech processing
  • Machine learning
  • Human-computer interaction

Background:

  • Speaker age and gender recognition is crucial for human-computer interaction (HCI) applications.
  • Current methods struggle with extracting salient speech features for accurate classification.
  • Challenges exist in developing robust models for age and gender identification from speech.

Purpose of the Study:

  • To propose a novel end-to-end convolutional neural network (CNN) model for speaker age and gender recognition.
  • To introduce a multi-attention module (MAM) for effective extraction of spatial and temporal speech features.
  • To enhance the accuracy and robustness of age and gender classification from speech signals.

Main Methods:

  • Developed a CNN model integrated with a specialized Multi-Attention Module (MAM).
  • MAM utilizes rectangular filters and separate time and frequency attention mechanisms.
  • Time attention focuses on temporal cues, while frequency attention targets relevant spatial frequency features.

Main Results:

  • Achieved 96% gender, 73% age, and 76% age-gender accuracy on the Common Voice dataset.
  • Attained 97% gender, 97% age, and 90% age-gender accuracy on a Korean speech dataset.
  • Demonstrated superior and robust performance in age, gender, and age-gender recognition tasks.

Conclusions:

  • The proposed CNN with MAM effectively extracts salient spatial and temporal features for speaker recognition.
  • The model shows high accuracy and robustness across different datasets for age and gender classification.
  • This approach advances the capabilities of speech processing in human-computer interaction.