Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Dynamic-based representation inconsistency and implicit constraints for offline reinforcement learning.

Neural networks : the official journal of the International Neural Network Society·2026
Same author

Prevalence of Kinesiophobia in Patients with Inflammatory Bowel Disease: A Cross-Sectional Study.

Journal of multidisciplinary healthcare·2026
Same author

<i>CENPF</i> Promotes Gastric Cancer Proliferation through c-Myc-Mediated GLS1 Upregulation and Glutamine Metabolism.

Oncology research·2026
Same author

Temporal dynamics and functional annotation of transcriptome rhythmicity in HEK293T cells.

PloS one·2026
Same author

Cobalt modulated local electronic environment of NiCoMnCuFe high-entropy alloys for stable overall water splitting.

Chemical communications (Cambridge, England)·2026
Same author

Organ-system-based psychiatry education in undergraduate medicine: learning outcomes and specialty choice.

BMC medical education·2026

Related Experiment Video

Updated: Nov 15, 2025

Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention
06:37

Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention

Published on: December 15, 2023

4.7K

AR3D: Attention Residual 3D Network for Human Action Recognition.

Min Dong1,2, Zhenglin Fang1, Yongfa Li1

  • 1School of Computer Science and Engineering, South China University of Technology, Guangzhou 510006, China.

Sensors (Basel, Switzerland)
|March 6, 2021
PubMed
Summary

This study introduces improved 3D CNN models, Residual 3D Network (R3D) and Attention Residual 3D Network (AR3D), for human action recognition. These models enhance temporal and global feature extraction, improving recognition performance.

Keywords:
3Daction recognitionattention mechanismconvolutional neural networkresidual

More Related Videos

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
03:31

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications

Published on: December 15, 2023

770
A Step-by-Step Implementation of DeepBehavior, Deep Learning Toolbox for Automated Behavior Analysis
05:41

A Step-by-Step Implementation of DeepBehavior, Deep Learning Toolbox for Automated Behavior Analysis

Published on: February 6, 2020

9.7K

Related Experiment Videos

Last Updated: Nov 15, 2025

Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention
06:37

Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention

Published on: December 15, 2023

4.7K
Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
03:31

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications

Published on: December 15, 2023

770
A Step-by-Step Implementation of DeepBehavior, Deep Learning Toolbox for Automated Behavior Analysis
05:41

A Step-by-Step Implementation of DeepBehavior, Deep Learning Toolbox for Automated Behavior Analysis

Published on: February 6, 2020

9.7K

Area of Science:

  • Computer Vision
  • Artificial Intelligence
  • Machine Learning

Background:

  • 2D CNNs independently process temporal and spatial features, limiting human action recognition.
  • 3D CNNs capture spatio-temporal features but suffer from high parameter counts and training difficulties.

Purpose of the Study:

  • To address limitations of existing deep learning models for video-based human action recognition.
  • To propose novel 3D CNN architectures that are efficient and effective.

Main Methods:

  • Developed a shallow feature extraction module to enhance temporal feature extraction in 3D CNNs.
  • Integrated a 3D spatio-temporal attention mechanism to improve global feature extraction.
  • Proposed Residual 3D Network (R3D) and Attention Residual 3D Network (AR3D) models, including fusion strategies (AR3D_V1, AR3D_V2).

Main Results:

  • The improved 3D residual structure reduces parameters and strengthens temporal feature extraction.
  • The 3D spatio-temporal attention module enhances global feature extraction for human actions.
  • Experimental results demonstrate performance improvements with fused R3D and AR3D structures compared to single components.

Conclusions:

  • The proposed R3D and AR3D models offer improved performance in human action recognition.
  • Combining residual structures and attention mechanisms in 3D CNNs is effective for enhancing spatio-temporal feature learning.
  • The novel architectures provide a more efficient and robust solution for video-based action recognition tasks.