Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Masking and Demasking Agents01:19

Masking and Demasking Agents

2.7K
EDTA titrations may necessitate masking and demasking agents to temporarily protect a particular metal ion in a mixture from the EDTA reaction. These agents facilitate the sequential analysis of the metal ions by forming stable complexes with some—but not all—metal ions during certain steps.
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
2.7K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Macrophage extracellular traps accelerate atherosclerosis progression via Rap1 pathway-mediated necroptosis and phenotypic switching in vascular smooth muscle cells.

Journal of translational medicine·2026
Same author

Metformin sensitizes esophageal squamous cell carcinoma to Vγ9Vδ2 T cell-mediated cytotoxicity by upregulating BTN3A1 and BTN2A1.

Cell death & disease·2026
Same author

The combined use of HA380 hemoperfusion in cardiopulmonary bypass alleviates postoperative inflammatory response and organ dysfunction following cardiac surgery.

Journal of cardiothoracic surgery·2026
Same author

Laser-Tracker-Based Robot Pose Measurement Using PSD Spot Sensing and Multi-Sensor Fusion with Simulation Validation.

Micromachines·2026
Same author

Differences in immune indicators among normal, high-risk, and esophageal cancer populations and development of a predictive model.

Frontiers in immunology·2026
Same author

Development of a cold-heat syndrome classification model for children with allergic rhinitis based on multimodal data.

Translational pediatrics·2026

Related Experiment Video

Updated: Sep 17, 2025

Author Spotlight: Revolutionizing Remote Surgery with Augmented Reality and Robotics for Enhanced Precision and Accessibility
07:46

Author Spotlight: Revolutionizing Remote Surgery with Augmented Reality and Robotics for Enhanced Precision and Accessibility

Published on: August 9, 2024

854

Local pattern aware 3D video swin transformer with masked autoencoding for realtime augmented reality gesture

Suli Wang1,2

  • 1Faculty of Data Science, City University of Macau, Taipa, 999078, Macau, China. wangsl@gcu.edu.cn.

Scientific Reports
|July 2, 2025
PubMed
Summary

This study introduces a real-time augmented reality gesture recognition algorithm using Swin Transformer and masked autoencoder, improving spatio-temporal feature extraction and performance. The novel approach enhances accuracy and efficiency for interactive applications.

Keywords:
3D video Swin transformerGesture recognitionLocal pattern awarenessMasked autoencodingReal-Time augmented reality

More Related Videos

Author Spotlight: Insights into the Analysis of Human Interaction with 3D Virtual Objects
06:36

Author Spotlight: Insights into the Analysis of Human Interaction with 3D Virtual Objects

Published on: October 18, 2024

1.1K
Photorealistic Learned Landscapes for Augmented Reality
06:54

Photorealistic Learned Landscapes for Augmented Reality

Published on: June 27, 2025

188

Related Experiment Videos

Last Updated: Sep 17, 2025

Author Spotlight: Revolutionizing Remote Surgery with Augmented Reality and Robotics for Enhanced Precision and Accessibility
07:46

Author Spotlight: Revolutionizing Remote Surgery with Augmented Reality and Robotics for Enhanced Precision and Accessibility

Published on: August 9, 2024

854
Author Spotlight: Insights into the Analysis of Human Interaction with 3D Virtual Objects
06:36

Author Spotlight: Insights into the Analysis of Human Interaction with 3D Virtual Objects

Published on: October 18, 2024

1.1K
Photorealistic Learned Landscapes for Augmented Reality
06:54

Photorealistic Learned Landscapes for Augmented Reality

Published on: June 27, 2025

188

Area of Science:

  • Computer Vision
  • Human-Computer Interaction
  • Machine Learning

Background:

  • Traditional Transformer models face challenges in spatio-temporal feature extraction and real-time performance for gesture interaction.
  • Efficient and accurate gesture recognition is crucial for advancing augmented reality applications.

Purpose of the Study:

  • To propose a real-time augmented reality gesture interaction algorithm leveraging Swin Transformer and masked autoencoder.
  • To enhance spatio-temporal feature extraction, real-time performance, and overall accuracy in gesture recognition.

Main Methods:

  • Utilized synthetic data generation for efficient 3D gesture annotation.
  • Developed an image denoising model using maximum a posteriori probability with weighted Euclidean distance and structural similarity optimization.
  • Integrated EfficientNet and Transformer models with skip connections and triplet attention for gesture detection and segmentation.
  • Introduced local texture feature prior (RTHLBP) for improved accuracy.
  • Employed a Vision Transformer (ViT) architecture with a masked autoencoder, dynamic weight fusion, and relative total variation map for gesture classification.
  • Implemented a multi-core parallel computing strategy for real-time performance optimization.

Main Results:

  • The proposed model achieved superior accuracy, F1 score, and mean intersection over union (MIoU) on the 4 GTEA sub-dataset compared to CNN, Transformer, MobileNet, and DenseNet models.
  • Demonstrated significant improvements in performance, especially on smaller datasets.
  • Showcased substantial reductions in computation time and maintained high computational efficiency with increased DSP cores.

Conclusions:

  • The developed algorithm effectively addresses the limitations of traditional models in real-time augmented reality gesture interaction.
  • The combination of advanced deep learning techniques and optimization strategies leads to state-of-the-art performance in gesture recognition and segmentation.
  • The optimized real-time performance makes the algorithm suitable for practical augmented reality applications.