Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Non-Verbal Cues01:29

Non-Verbal Cues

213
Non-verbal communication extends beyond gestures and facial expressions to include vocal elements known as paralanguage. Paralanguage consists of non-verbal vocal cues such as pitch, loudness, speech rate, pauses, and non-verbal vocalizations like laughter, sighs, and moans. These elements not only accompany speech but also provide critical emotional and contextual information.The Role of Paralanguage in CommunicationParalanguage adds depth to spoken language by conveying emotions and...
213
Labeling Emotion01:20

Labeling Emotion

560
Emotional labeling is a cognitive process that involves identifying and naming one's emotions, such as anger, fear, happiness, or sadness. It allows individuals to recognize and express their internal emotional states, a critical aspect of emotional regulation and communication. Labeling emotions requires more than mere recognition; it also involves drawing upon memory and contextual cues to understand the current situation and apply a corresponding emotional label. For instance, feeling...
560
Physiology of Emotion01:20

Physiology of Emotion

3.0K
The physiology of emotions is a multifaceted process involving the autonomic nervous system, brain structures, hormones, and neurotransmitters. This intricate interplay dictates how emotions manifest in the body and influence behavior.
Autonomic Nervous System
The autonomic nervous system (ANS) plays a critical role in emotional responses by regulating involuntary physiological functions. It consists of two main components: the sympathetic and parasympathetic systems. The sympathetic system...
3.0K
Facial Feedback Hypothesis01:24

Facial Feedback Hypothesis

493
Charles Darwin proposed that facial expressions are an evolutionary adaptation for communication. He argued that these expressions are not influenced by culture but are universal across species. For example, a snarling expression with exposed teeth signals a threat in many animals, including humans. Darwin also suggested that displaying an emotion can intensify the feeling. Smiling, for example, could enhance one's sense of happiness. This idea laid the foundation for understanding the role...
493
Classification of Signals01:30

Classification of Signals

1.3K
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
1.3K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Age and Gender Recognition Using a Convolutional Neural Network with a Specially Designed Multi-Attention Module through Speech Spectrograms.

Sensors (Basel, Switzerland)·2021
Same author

Deep-Net: A Lightweight CNN-Based Speech Emotion Recognition System Using Deep Frequency Features.

Sensors (Basel, Switzerland)·2020
See all related articles

Related Experiment Video

Updated: Dec 31, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.9K

A CNN-Assisted Enhanced Audio Signal Processing for Speech Emotion Recognition.

Mustaqeem1, Soonil Kwon1

  • 1Interaction Technology Laboratory, Department of Software, Sejong University, Seoul 05006, Korea.

Sensors (Basel, Switzerland)
|January 8, 2020
PubMed
Summary

This study introduces an AI-assisted deep stride convolutional neural network (DSCNN) for accurate speech emotion recognition (SER). The model enhances accuracy while reducing computational complexity for real-world applications.

Keywords:
artificial intelligenceemotion recognitionneural networksnoise removalsignals enhancementspectrogram

More Related Videos

Asthma Detection Research Based on Voice Signal Processing and Machine Learning
04:04

Asthma Detection Research Based on Voice Signal Processing and Machine Learning

Published on: July 22, 2025

835
Exploring the Use of Isolated Expressions and Film Clips to Evaluate Emotion Recognition by People with Traumatic Brain Injury
05:51

Exploring the Use of Isolated Expressions and Film Clips to Evaluate Emotion Recognition by People with Traumatic Brain Injury

Published on: May 15, 2016

9.4K

Related Experiment Videos

Last Updated: Dec 31, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.9K
Asthma Detection Research Based on Voice Signal Processing and Machine Learning
04:04

Asthma Detection Research Based on Voice Signal Processing and Machine Learning

Published on: July 22, 2025

835
Exploring the Use of Isolated Expressions and Film Clips to Evaluate Emotion Recognition by People with Traumatic Brain Injury
05:51

Exploring the Use of Isolated Expressions and Film Clips to Evaluate Emotion Recognition by People with Traumatic Brain Injury

Published on: May 15, 2016

9.4K

Area of Science:

  • Artificial Intelligence
  • Human-Computer Interaction
  • Speech Signal Processing

Background:

  • Speech is a primary communication method and a key area for human-computer interaction (HCI).
  • Speech emotion recognition (SER) is crucial for applications like healthcare and virtual reality, aiming to identify a speaker's emotional state.
  • Current SER models face challenges in accuracy and computational efficiency.

Purpose of the Study:

  • To improve the accuracy of speech emotion recognition (SER) beyond current state-of-the-art methods.
  • To decrease the computational complexity of the proposed SER model.
  • To develop an effective SER model applicable to real-world scenarios.

Main Methods:

  • Proposed an artificial intelligence-assisted deep stride convolutional neural network (DSCNN) architecture.
  • Utilized a plain nets strategy to learn salient features from speech spectrograms.
  • Employed convolutional layers with strides for down-sampling and fully connected layers for global feature learning, followed by a SoftMax classifier.

Main Results:

  • Achieved accuracy improvements of 7.85% on the IEMOCAP dataset and 4.5% on the RAVDESS dataset.
  • Reduced the model size by 34.5 MB, indicating decreased computational complexity.
  • Demonstrated superior performance compared to existing state-of-the-art SER techniques.

Conclusions:

  • The proposed DSCNN model significantly enhances SER accuracy and reduces computational load.
  • The technique proves effective and shows strong potential for real-world applications in HCI and beyond.
  • This research contributes to advancing the field of emotion recognition from speech signals.