Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Vision01:24

Vision

Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Multiaxial Fatigue Life Assessment of Large Welded Flange Shafts: A Continuum Damage Mechanics Approach.

Materials (Basel, Switzerland)·2025
Same author

GR-AttNet: Robotic grasping with lightweight spatial attention mechanism.

PloS one·2025
Same author

Dual-branch differential channel hypergraph convolutional network for human skeleton based action recognition.

PloS one·2025
Same author

MEP-YOLOv5s: Small-Target Detection Model for Unmanned Aerial Vehicle-Captured Images.

Sensors (Basel, Switzerland)·2025
Same author

Highly active and reversible NiPSe<sub>3</sub> anode for sodium-ion batteries: enabling ultrafast sodium storage with exceptional cycling stability.

Journal of colloid and interface science·2025
Same author

Use of BOIvy Optimization Algorithm-Based Machine Learning Models in Predicting the Compressive Strength of Bentonite Plastic Concrete.

Materials (Basel, Switzerland)·2025

Related Experiment Video

Updated: Jul 25, 2026

A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
04:23

A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images

Published on: April 21, 2023

1.7K

LS-VIT: Vision Transformer for action recognition based on long and short-term temporal difference.

Dong Chen1,2,3, Peisong Wu1,3, Mingdong Chen1,3

  • 1College of Physics and Electronic Engineering, Nanning Normal University, Nanning, China.

Frontiers in Neurorobotics
|November 15, 2024
PubMed
Summary

This study introduces the Long and Short-term Temporal Difference Vision Transformer (LS-VIT) for efficient 3D video action recognition. The LS-VIT model achieves high accuracy by effectively capturing both short-term and long-term motion details in videos.

Keywords:
Vision Transformeraction recognitiondeep learningmotion extractiontemporal crossing fusion

More Related Videos

Author Spotlight: Revolutionizing Remote Surgery with Augmented Reality and Robotics for Enhanced Precision and Accessibility
07:46

Author Spotlight: Revolutionizing Remote Surgery with Augmented Reality and Robotics for Enhanced Precision and Accessibility

Published on: August 9, 2024

651
Author Spotlight: Deciphering Electrical Networks Behind Complex Brain Activities and Disorders
05:49

Author Spotlight: Deciphering Electrical Networks Behind Complex Brain Activities and Disorders

Published on: November 1, 2024

717

Related Experiment Videos

Last Updated: Jul 25, 2026

A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
04:23

A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images

Published on: April 21, 2023

1.7K
Author Spotlight: Revolutionizing Remote Surgery with Augmented Reality and Robotics for Enhanced Precision and Accessibility
07:46

Author Spotlight: Revolutionizing Remote Surgery with Augmented Reality and Robotics for Enhanced Precision and Accessibility

Published on: August 9, 2024

651
Author Spotlight: Deciphering Electrical Networks Behind Complex Brain Activities and Disorders
05:49

Author Spotlight: Deciphering Electrical Networks Behind Complex Brain Activities and Disorders

Published on: November 1, 2024

717

Area of Science:

  • Computer Vision
  • Artificial Intelligence
  • Machine Learning

Background:

  • Transformer models excel in 2D vision but face computational challenges in 3D video tasks like action recognition.
  • Directly applying temporal transformations to 3D video data increases computational and memory demands due to data patch multiplication and complex self-attention mechanisms.

Purpose of the Study:

  • To develop an efficient and precise 3D self-attentive model for video action recognition.
  • To address the computational challenges posed by transformer models in 3D video analysis.

Main Methods:

  • Introduced the Long and Short-term Temporal Difference Vision Transformer (LS-VIT).
  • Incorporated short-term motion details by weighting differences across consecutive frames.
  • Integrated a module for long-term motion understanding using motion excitation and temporal differences from various segments.

Main Results:

  • LS-VIT achieved high recognition accuracy on multiple benchmarks, including UCF101, HMDB51, and Kinetics-400.
  • The model effectively models both short-term and long-term motion dynamics in videos.

Conclusions:

  • LS-VIT demonstrates strong performance in 3D video action recognition.
  • The model shows potential for further optimization to enhance real-time performance and action prediction capabilities.