Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Behavioral Interpretation of Willingness to Use Wearable Health Devices in Community Residents: A Cross-Sectional Study.

International journal of environmental research and public health·2023
Same author

Two Stable Sodalite-Cage-Based MOFs for Highly Gas Selective Capture and Conversion in Cycloaddition Reaction.

ACS applied materials & interfaces·2023
Same author

Nonstructural protein 2A2 from Duck hepatitis A virus type 1 inhibits interferon beta production by interaction with mitochondrial antiviral signaling protein and TANK-binding kinase 1.

Veterinary microbiology·2023
Same author

Glutathione peroxidase-like nanozymes: mechanism, classification, and bioapplication.

Biomaterials science·2023
Same author

IFN-γ-STAT1-mediated CD8<sup>+</sup> T-cell-neural stem cell cross talk controls astrogliogenesis after spinal cord injury.

Inflammation and regeneration·2023
Same author

Expression and Localization of Fas-Associated Factor 1 in Testicular Tissues of Different Ages and Ovaries at Different Reproductive Cycle Phases of <i>Bos grunniens</i>.

Animals : an open access journal from MDPI·2023

Related Experiment Video

Updated: Jul 26, 2025

A Methodology for Capturing Joint Visual Attention Using Mobile Eye-Trackers
12:39

A Methodology for Capturing Joint Visual Attention Using Mobile Eye-Trackers

Published on: January 18, 2020

7.7K

Learning Cross-Attention Discriminators via Alternating Time-Space Transformers for Visual Tracking.

Wuwei Wang, Ke Zhang, Yu Su

    IEEE Transactions on Neural Networks and Learning Systems
    |June 20, 2023
    PubMed
    Summary

    This study introduces a pure Transformer-based visual tracking model, the alternating time-space Transformers (ATSTs), which excels by using attention mechanisms instead of convolution. ATSTs achieve competitive performance with less training data, outperforming convolutional trackers.

    More Related Videos

    Exploring Infant Sensitivity to Visual Language using Eye Tracking and the Preferential Looking Paradigm
    06:07

    Exploring Infant Sensitivity to Visual Language using Eye Tracking and the Preferential Looking Paradigm

    Published on: May 15, 2019

    8.4K
    Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
    08:25

    Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment

    Published on: May 7, 2019

    9.0K

    Related Experiment Videos

    Last Updated: Jul 26, 2025

    A Methodology for Capturing Joint Visual Attention Using Mobile Eye-Trackers
    12:39

    A Methodology for Capturing Joint Visual Attention Using Mobile Eye-Trackers

    Published on: January 18, 2020

    7.7K
    Exploring Infant Sensitivity to Visual Language using Eye Tracking and the Preferential Looking Paradigm
    06:07

    Exploring Infant Sensitivity to Visual Language using Eye Tracking and the Preferential Looking Paradigm

    Published on: May 15, 2019

    8.4K
    Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
    08:25

    Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment

    Published on: May 7, 2019

    9.0K

    Area of Science:

    • Computer Vision
    • Artificial Intelligence
    • Machine Learning

    Background:

    • Convolutional Neural Networks (CNNs) dominate visual tracking but struggle with long-range spatial dependencies.
    • Transformer-assisted methods combine CNNs and Transformers, improving feature representation.
    • Existing approaches often integrate Transformers with CNNs, limiting pure Transformer exploration.

    Purpose of the Study:

    • To propose a novel, pure Transformer-based visual tracking model.
    • To overcome the limitations of convolutional operations in relating spatially distant information.
    • To introduce a semi-Siamese architecture leveraging attention mechanisms exclusively.

    Main Methods:

    • Developed multistage alternating time-space Transformers (ATSTs) inspired by Vision Transformers (ViTs).
    • Employed a time-space self-attention module for feature extraction backbone.
    • Utilized a cross-attention discriminator for direct response map estimation without convolution or correlation filters.

    Main Results:

    • The ATST-based model demonstrated favorable results against state-of-the-art convolutional trackers.
    • Achieved comparable performance to recent "CNN + Transformer" trackers on multiple benchmarks.
    • Required significantly less training data compared to existing methods.

    Conclusions:

    • Pure Transformer architectures are viable and effective for visual tracking.
    • The proposed ATST model offers a promising direction for efficient and robust visual tracking.
    • Attention-based mechanisms can replace convolution for enhanced feature representation in tracking.