Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Alternating Current-Driven Diol Epimerization via a Deplete-Regenerate Strategy.

Journal of the American Chemical Society·2025
Same author

Correlation Information Enhanced Graph Anomaly Detection via Hypergraph Transformation.

IEEE transactions on cybernetics·2025
Same author

Iron-Catalyzed Reductive Allylic C─H Amination of Olefin with Nitroarenes via Intermolecular Nitroso Ene Reaction.

Angewandte Chemie (International ed. in English)·2025
Same author

Boundary-enhanced local-global collaborative network for medical image segmentation.

Scientific reports·2025
Same author

Implementing a Social Presence-Based Teaching Strategy in Online Lecture Learning.

European journal of investigation in health, psychology and education·2024
Same author

Flow2GNN: Flexible Two-Way Flow Message Passing for Enhancing GNNs Beyond Homophily.

IEEE transactions on cybernetics·2024

Related Experiment Video

Updated: Aug 13, 2025

Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention
06:37

Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention

Published on: December 15, 2023

4.0K

Emotion Recognition from Large-Scale Video Clips with Cross-Attention and Hybrid Feature Weighting Neural Networks.

Siwei Zhou1, Xuemei Wu1, Fan Jiang1

  • 1Key Laboratory of Intelligent Education Technology and Application of Zhejiang Province, Zhejiang Normal University, Jinhua 321004, China.

International Journal of Environmental Research and Public Health
|January 21, 2023
PubMed
Summary

This study introduces a novel network for accurate emotion recognition from videos, improving upon existing methods by effectively fusing facial and contextual information. The model enhances understanding of human emotions for applications like mental health and job stress assessment.

Keywords:
attention mechanismcross-channeldeep convolutional neural networkdeep feature fusionemotion recognitionlarge-scale video clips

More Related Videos

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
03:31

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications

Published on: December 15, 2023

600
Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
08:25

Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment

Published on: May 7, 2019

9.1K

Related Experiment Videos

Last Updated: Aug 13, 2025

Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention
06:37

Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention

Published on: December 15, 2023

4.0K
Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
03:31

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications

Published on: December 15, 2023

600
Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
08:25

Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment

Published on: May 7, 2019

9.1K

Area of Science:

  • Computer Vision
  • Affective Computing
  • Machine Learning

Background:

  • Human emotion recognition is vital for understanding mental states and has applications in mental health, job stress, and tourist satisfaction.
  • Computer vision techniques are widely used for emotion recognition from visual media, but existing models often fail to effectively integrate facial and contextual features.
  • This limitation leads to emotion confusion and misunderstanding, highlighting the need for improved feature fusion methods.

Purpose of the Study:

  • To develop a novel cross-attention and hybrid feature weighting network for accurate emotion recognition from large-scale video clips.
  • To effectively exploit the complementary information between face and context features for enhanced emotion detection.
  • To address the limitations of existing models in capturing inter-feature interactions and feature fusion.

Main Methods:

  • A dual-branch encoding (DBE) network generates initial face and context features.
  • A hierarchical-attention encoding (HAE) network employs cross-attention (CA) for feature complementarity and element recalibration (ER) for feature map revision.
  • An adaptive-attention (AA) block infers optimal feature fusion weights, followed by a deep fusion (DF) block for final emotion state prediction.

Main Results:

  • The proposed model demonstrated effectiveness in emotion recognition tasks on the CAER-S dataset.
  • The cross-attention and hybrid feature weighting approach successfully captured complementary information between facial and contextual cues.
  • Experimental results validated the model's ability to overcome emotion confusion and misunderstanding.

Conclusions:

  • The novel network effectively integrates facial and contextual information for superior emotion recognition in videos.
  • The method shows significant potential for real-world applications such as analyzing tourist reviews, estimating job stress, and assessing mental health.
  • This research advances the field of affective computing by providing a more robust approach to understanding human emotions from visual data.