Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Video

Updated: Sep 13, 2025

Capturing Dynamic Finger Gesturing with High-resolution Surface Electromyography and Computer Vision
08:15

Capturing Dynamic Finger Gesturing with High-resolution Surface Electromyography and Computer Vision

Published on: March 28, 2025

746

Fusion of Multimodal Spatio-Temporal Features and 3D Deformable Convolution Based on Sign Language Recognition in

Qian Zhou1, Hui Li1, Weizhi Meng2

  • 1School of Computer Science, Nanjing University of Posts and Telecommunications, 9 Wenyuan Road, Nanjing 210023, China.

Sensors (Basel, Switzerland)
|July 30, 2025
PubMed
Summary

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Y-DWMS: A Digital Watermark Management System Based on Smart Contracts.

Sensors (Basel, Switzerland)·2019
Same author

[Association between genetic polymorphism of tumor necrosis factor and chronic severe hepatitis B in patients].

Zhonghua yi xue za zhi·2007
Same author

In vivo translational inaccuracy in Escherichia coli: missense reporting using extremely low activity mutants of Vibrio harveyi luciferase.

Biochemistry·2007
Same author

[Construction of recombinant adenovirus vector expressing extracellular domain of TbetaR-II-RANTES fusion gene and its anti-tumor effects].

Zhonghua zhong liu za zhi [Chinese journal of oncology]·2007
Same author

[Characteristics, evolution and variation of M genes of human avian H5N1 strains in Guangdong].

Bing du xue bao = Chinese journal of virology·2007
Same author

Dynamic changes in microbial activity and community structure during biodegradation of petroleum compounds: a laboratory experiment.

Journal of environmental sciences (China)·2007

This study introduces a novel deep learning framework for accurate sign language recognition (SLR) using skeleton and image data. The multimodal approach achieves competitive results on public datasets, advancing gesture interpretation.

Area of Science:

  • Computer Vision
  • Artificial Intelligence
  • Human-Computer Interaction

Background:

  • Sign language recognition (SLR) is challenging due to the complex spatio-temporal dynamics of visual language.
  • Existing methods often struggle with precise and efficient interpretation of raw video data.

Purpose of the Study:

  • To develop a novel multimodal deep learning framework for accurate and efficient sign language recognition.
  • To effectively integrate information from skeleton data and raw RGB images for improved SLR performance.

Main Methods:

  • A Multi-Stream Spatio-Temporal Graph Convolutional Network (MSGCN) was proposed for skeleton feature extraction, incorporating decoupling graph convolution, self-emphasizing temporal convolution, and spatio-temporal joint attention.
  • A 3D ResNet model with deformable convolution (D-ResNet) was utilized to process raw RGB image sequences.
Keywords:
ResNet2+1Dmultimodal fusionsign language recognitionspatio-temporal graph convolutional network

More Related Videos

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
03:31

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications

Published on: December 15, 2023

637

Related Experiment Videos

Last Updated: Sep 13, 2025

Capturing Dynamic Finger Gesturing with High-resolution Surface Electromyography and Computer Vision
08:15

Capturing Dynamic Finger Gesturing with High-resolution Surface Electromyography and Computer Vision

Published on: March 28, 2025

746
Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
03:31

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications

Published on: December 15, 2023

637
  • A Multi-Stream Fusion Module (MFM) with a gating mechanism was employed to merge features from both modalities.
  • Main Results:

    • The proposed multimodal framework achieved competitive performance on the AUTSL and WLASL public datasets.
    • The integration of skeleton and RGB modalities through the MSGCN and D-ResNet models, fused by MFM, demonstrated effectiveness in SLR.

    Conclusions:

    • The novel multimodal deep learning approach offers a promising direction for advancing sign language recognition technology.
    • The MSGCN and D-ResNet models, combined with the MFM, provide a robust solution for capturing complex spatio-temporal information in sign language gestures.