Related Experiment Video
Updated: Sep 9, 2025

Using Looming Visual Stimuli to Evaluate Mouse Vision
Published on: June 13, 2019
Efficient High-Order Spatial Interactions for Visual Perception
Researchers developed Recursive Gated Convolution (g nConv) to efficiently implement key vision Transformer features using convolutions. This new operation enhances various vision models, improving performance across image recognition, 3D analysis, and vision-language tasks.
Area of Science:
- Computer Vision
- Deep Learning
- Artificial Intelligence
Background:
- Vision Transformers (ViTs) achieve success via self-attention's spatial modeling.
- Convolutional Neural Networks (CNNs) are foundational in computer vision.
- Integrating ViT strengths into CNNs is an active research area.
Purpose of the Study:
- To introduce a convolution-based framework that replicates ViT's spatial modeling.
- To develop a novel operation, Recursive Gated Convolution (g nConv), for high-order spatial interactions.
- To create versatile backbones (HorNet, Hor3D, HorCLIP) for diverse visual tasks.
Main Methods:
- Proposed Recursive Gated Convolution (g nConv) for efficient, high-order spatial interactions.
- Developed generic vision backbones: HorNet (image recognition), Hor3D (point clouds), HorCLIP (vision-language).
- Integrated g nConv as a plug-and-play module into existing architectures.
Main Results:
- HorNet outperforms Swin Transformers and ConvNeXt on ImageNet, COCO, and ADE20K.
- g nConv improves dense prediction tasks with reduced computation.
- Hor3D shows efficacy in 3D semantic segmentation; HorCLIP excels in vision-language tasks.
Conclusions:
- g nConv effectively combines ViT and CNN merits, offering a new basic operation for visual modeling.
- The proposed HorNet family demonstrates strong performance and scalability.
- High-order spatial interactions via g nConv are beneficial across various visual modalities and tasks.
More Related Videos
07:45Assessing Binocular Central Visual Field and Binocular Eye Movements in a Dichoptic Viewing Condition
Published on: July 21, 2020
07:12Development of a Gaze-Contingent Display Framework Designed for Perceptual and Oculomotor Research with Simulated Central Vision Loss
Published on: April 11, 2025
Related Concept Videos
Depth Perception and Spatial Vision
Visual System
Once through the pupil, the light passes through the lens, a...
Vision
Parallel Processing
Anatomy of the Eyeball
Gestalt Principles of Perception