Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Vision01:24

Vision

52.9K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
52.9K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

CoNav: Collaborative Cross-Modal Reasoning for Embodied Navigation.

IEEE transactions on pattern analysis and machine intelligence·2026
Same author

Role of NADPH oxidase 2-derived reactive oxygen species in cardiac electrophysiological disorders.

Channels (Austin, Tex.)·2026
Same author

Cyber Defense Effectiveness Evaluation for ICS Under Uncertainty: A Dynamic Bayesian Network Approach with Information Entropy.

Entropy (Basel, Switzerland)·2026
Same author

vNOTES Pelvic Reconstruction and Presacral-Uterosacral Ligament Compound Suspension for Treatment of Multicompartment Pelvic Organ Prolapse: A Three-Year, Three-Arm, Open-Label, Randomized Controlled Trial.

Journal of minimally invasive gynecology·2026
Same author

Bridging Subjectivity in Affective Explanation Captioning via Consensus-Prompted Emotion Reasoning.

IEEE transactions on image processing : a publication of the IEEE Signal Processing Society·2026
Same author

Looking Broader for Knowledge Distillation Via Receptive-Field Alignment.

IEEE transactions on pattern analysis and machine intelligence·2026

Related Experiment Video

Updated: May 24, 2025

A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
04:23

A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images

Published on: April 21, 2023

1.7K

DHVT: Dynamic Hybrid Vision Transformer for Small Dataset Recognition.

Zhiying Lu, Chuanbin Liu, Xiaojun Chang

    IEEE Transactions on Pattern Analysis and Machine Intelligence
    |March 3, 2025
    PubMed
    Summary

    This study introduces the Dynamic Hybrid Vision Transformer (DHVT) and DHVT2, enhancing Vision Transformers (ViTs) for better image recognition. These models bridge the performance gap with Convolutional Neural Networks (CNNs), especially on limited datasets.

    More Related Videos

    Author Spotlight: Assessment of Visual Acuity in Central Vision Loss Through Motion-Based Peripheral Vision Testing
    06:25

    Author Spotlight: Assessment of Visual Acuity in Central Vision Loss Through Motion-Based Peripheral Vision Testing

    Published on: February 23, 2024

    515
    Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique
    04:48

    Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique

    Published on: July 5, 2024

    348

    Related Experiment Videos

    Last Updated: May 24, 2025

    A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
    04:23

    A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images

    Published on: April 21, 2023

    1.7K
    Author Spotlight: Assessment of Visual Acuity in Central Vision Loss Through Motion-Based Peripheral Vision Testing
    06:25

    Author Spotlight: Assessment of Visual Acuity in Central Vision Loss Through Motion-Based Peripheral Vision Testing

    Published on: February 23, 2024

    515
    Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique
    04:48

    Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique

    Published on: July 5, 2024

    348

    Area of Science:

    • Computer Science
    • Artificial Intelligence
    • Machine Learning

    Background:

    • Vision Transformers (ViTs) exhibit a performance gap compared to Convolutional Neural Networks (CNNs) when trained on limited datasets due to insufficient inductive bias.
    • Key limitations in ViTs include inadequate spatial relevance and suboptimal channel representation, hindering their ability to capture fine-grained features and robust data patterns.

    Purpose of the Study:

    • To address the limitations of ViTs by proposing the Dynamic Hybrid Vision Transformer (DHVT) and its computationally efficient variant, DHVT2.
    • To enhance spatial feature extraction and improve channel representation within the Vision Transformer architecture.

    Main Methods:

    • DHVT integrates convolutional operations into the feature embedding and projection phases to bolster spatial relevance.
    • A dynamic aggregation mechanism and a novel 'head token' are employed to recalibrate and harmonize channel representations.
    • The study explores optimal network meta-structures, adopting a multi-stage hybrid design without a class token, further refined by a dimensional variable residual connection mechanism in DHVT2.

    Main Results:

    • DHVT and DHVT2 achieve state-of-the-art results in image recognition tasks, effectively narrowing the performance disparity between ViTs and CNNs.
    • The proposed models demonstrate strong generalization capabilities in downstream experiments.

    Conclusions:

    • The Dynamic Hybrid Vision Transformer (DHVT) and DHVT2 successfully overcome the inherent limitations of standard ViTs, particularly in data-scarce scenarios.
    • These hybrid architectures offer a promising direction for advancing vision-based artificial intelligence, providing competitive performance and enhanced generalization.