Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Labeling Emotion01:20

Labeling Emotion

475
Emotional labeling is a cognitive process that involves identifying and naming one's emotions, such as anger, fear, happiness, or sadness. It allows individuals to recognize and express their internal emotional states, a critical aspect of emotional regulation and communication. Labeling emotions requires more than mere recognition; it also involves drawing upon memory and contextual cues to understand the current situation and apply a corresponding emotional label. For instance, feeling...
475
Non-Verbal Cues01:29

Non-Verbal Cues

142
Non-verbal communication extends beyond gestures and facial expressions to include vocal elements known as paralanguage. Paralanguage consists of non-verbal vocal cues such as pitch, loudness, speech rate, pauses, and non-verbal vocalizations like laughter, sighs, and moans. These elements not only accompany speech but also provide critical emotional and contextual information.The Role of Paralanguage in CommunicationParalanguage adds depth to spoken language by conveying emotions and...
142
Classification of Signals01:30

Classification of Signals

1.2K
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
1.2K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Preliminary results of a new endoscopic underlay cartilage tympanoplasty with lateral malleolar flap.

European archives of oto-rhino-laryngology : official journal of the European Federation of Oto-Rhino-Laryngological Societies (EUFOS) : affiliated with the German Society for Oto-Rhino-Laryngology - Head and Neck Surgery·2025
Same author

Deep learning for diagnosis of COVID-19 using 3D CT scans.

Computers in biology and medicine·2021
Same author

Gabor wavelet-based deep learning for skin lesion classification.

Computers in biology and medicine·2019
Same author

Performance analysis of different classification algorithms using different feature selection methods on Parkinson's disease detection.

Journal of neuroscience methods·2018
Same author

Effects of different covariates and contrasts on classification of Parkinson's disease using structural MRI.

Computers in biology and medicine·2018
Same author

The Role of Topical Thymoquinone in the Treatment of Acute Otitis Externa; an Experimental Study in Rats.

The journal of international advanced otology·2017

Related Experiment Video

Updated: Nov 27, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.8K

3D CNN-Based Speech Emotion Recognition Using K-Means Clustering and Spectrograms.

Noushin Hajarolasvadi1, Hasan Demirel1

  • 1Department of Electrical and Electronics Engineering, Eastern Mediterranean University, 99628 Gazimagusa, North Cyprus, via Mersin 10, Turkey.

Entropy (Basel, Switzerland)
|December 3, 2020
PubMed
Summary

This study introduces a novel speech-based emotion recognition system using 3D Convolutional Neural Networks (CNNs). The method effectively identifies emotions from audio features and spectrograms, outperforming existing state-of-the-art approaches.

Keywords:
3D convolutional neural networksdeep learningk-means clusteringspectrogramsspeech emotion recognition

More Related Videos

Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention
06:37

Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention

Published on: December 15, 2023

4.9K
Exploring the Use of Isolated Expressions and Film Clips to Evaluate Emotion Recognition by People with Traumatic Brain Injury
05:51

Exploring the Use of Isolated Expressions and Film Clips to Evaluate Emotion Recognition by People with Traumatic Brain Injury

Published on: May 15, 2016

9.3K

Related Experiment Videos

Last Updated: Nov 27, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.8K
Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention
06:37

Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention

Published on: December 15, 2023

4.9K
Exploring the Use of Isolated Expressions and Film Clips to Evaluate Emotion Recognition by People with Traumatic Brain Injury
05:51

Exploring the Use of Isolated Expressions and Film Clips to Evaluate Emotion Recognition by People with Traumatic Brain Injury

Published on: May 15, 2016

9.3K

Area of Science:

  • Artificial Intelligence
  • Speech Processing
  • Human-Computer Interaction

Background:

  • Emotion recognition is crucial for enhancing human-robot interaction.
  • Accurate emotion detection from speech remains a significant research challenge.

Purpose of the Study:

  • To propose an effective emotion recognition system utilizing speech signal analysis.
  • To improve the accuracy of emotion detection in human-robot interaction contexts.

Main Methods:

  • Speech signals are segmented into overlapping frames, and audio features (e.g., MFCC, pitch, intensity) are extracted.
  • Spectrograms are generated, and k-means clustering identifies keyframes to summarize signals.
  • A 3D Convolutional Neural Network (CNN) is trained on 3D tensors of keyframe spectrograms using 10-fold cross-validation.

Main Results:

  • The proposed 3D CNN model demonstrated superior performance compared to state-of-the-art methods.
  • Experiments were validated on diverse emotion databases: SAVEE, RML, and eNTERFACE'05.

Conclusions:

  • The developed speech-based emotion recognition system offers a promising approach for human-robot interaction.
  • The methodology effectively leverages audio features and deep learning for accurate emotion detection.