Related Experiment Video
Updated: Jun 30, 2025

06:25
Time Multiplexing Super Resolving Technique for Imaging from a Moving Platform
Published on: February 12, 2014
8.5K
HIRI-ViT: Scaling Vision Transformer With High Resolution Inputs.
Summary
We introduce HIRI-ViT, a novel hybrid deep learning backbone for computer vision tasks. This model efficiently processes high-resolution images, achieving state-of-the-art accuracy on image recognition and dense prediction benchmarks.
Area of Science:
- Computer Vision
- Deep Learning Architectures
- Image Recognition
Background:
- Hybrid deep models combining Vision Transformer (ViT) and Convolutional Neural Networks (CNNs) are effective for vision tasks.
- Increasing input resolution enhances model capacity but leads to computationally expensive, quadratically scaling costs.
- Existing hybrid backbones struggle with efficiency when handling high-resolution inputs.
Purpose of the Study:
- To develop a novel hybrid deep backbone, HIRI-ViT, specifically designed for efficient processing of high-resolution inputs.
- To improve computational efficiency and model performance in vision tasks by addressing the quadratic scaling issue of high-resolution inputs.
- To establish a new state-of-the-art in image recognition and dense prediction tasks with a cost-efficient architecture.
Main Methods:
- Proposed HIRI-ViT, a five-stage Vision Transformer (ViT) backbone tailored for high-resolution inputs.
- Implemented a novel approach decomposing CNN operations into two parallel branches: one high-resolution with fewer convolutions, and one low-resolution with more convolutions after down-sampling.
- Evaluated HIRI-ViT on ImageNet-1K (recognition) and COCO, ADE20K (dense prediction) datasets.
Main Results:
- HIRI-ViT demonstrated superior performance on both image recognition and dense prediction tasks.
- Achieved the best published Top-1 accuracy of 84.3% on ImageNet with 448x448 inputs, surpassing previous methods.
- Under comparable computational cost (~5.0 GFLOPs), HIRI-ViT significantly improved accuracy compared to existing models like iFormer-S.
Conclusions:
- HIRI-ViT offers a highly effective and computationally efficient solution for processing high-resolution images in deep learning.
- The proposed parallel CNN branch decomposition effectively balances feature resolution and computational cost.
- HIRI-ViT sets a new benchmark for performance in vision tasks, particularly for high-resolution image analysis.
More Related Videos
Related Concept Videos
Transformers with Off-Nominal Turns Ratios
152
In scenarios involving parallel transformers with disparate ratings, developing per-unit models requires accommodating off-nominal turns ratios. This situation arises when the selected base voltages are not proportional to the transformer’s voltage ratings. Consider a transformer where the rated voltages are related by the term a. If the chosen voltage bases satisfy a relationship involving term b, term c is defined as the ratio of these bases. This ratio is then substituted into the...
152
Types Of Transformers
974
Transformers can provide desired voltages to a circuit by modifying the number of turns in the secondary windings.
If the ratio of the number of turns in the secondary winding to that of the primary winding is greater than one, then the transformer is said to be a step-up transformer. In a step-up transformer, the voltage at the secondary winding is greater than the voltage applied at the primary winding.
However, if this ratio is less than one, the transformer is said to be a step-down...
If the ratio of the number of turns in the secondary winding to that of the primary winding is greater than one, then the transformer is said to be a step-up transformer. In a step-up transformer, the voltage at the secondary winding is greater than the voltage applied at the primary winding.
However, if this ratio is less than one, the transformer is said to be a step-down...
974
Reducing Line Loss
152
In a three-phase circuit, line loss is an indicator of energy dissipated as heat due to the resistance of transmission lines. To address this, incorporating transformers into the system—a step-up transformer at the source and a step-down transformer at the load—is a strategic solution. Two three-phase transformers are introduced to improve this.
With a step-up transformer at the source, the voltage is increased, thereby reducing the current in the transmission lines since power loss...
With a step-up transformer at the source, the voltage is increased, thereby reducing the current in the transmission lines since power loss...
152
Depth Perception and Spatial Vision
646
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
646
The Ideal Transformer
391
In single-phase two-winding transformers, two windings are coiled around a magnetic core characterized by cross-sectional area A and magnetic permeability μ. A phasor current i1 enters the left winding while i2 exits the right winding, establishing the fundamental working of the transformer through electromagnetic principles.
Ampere's Law forms the basis of understanding the magnetic field within the transformer. It states that the integral of the magnetic field intensity's...
Ampere's Law forms the basis of understanding the magnetic field within the transformer. It states that the integral of the magnetic field intensity's...
391
Transformers in Distribution System
102
Transformers in distribution systems can be broadly categorized into distribution substation transformers and other distribution transformers. They are crucial for stepping down high transmission voltages to levels suitable for distribution and end-user applications.
Distribution substation transformers come in various ratings and typically use mineral oil for insulation and cooling. To prevent moisture and air from entering the oil, some transformers use an inert gas like nitrogen to fill the...
Distribution substation transformers come in various ratings and typically use mineral oil for insulation and cooling. To prevent moisture and air from entering the oil, some transformers use an inert gas like nitrogen to fill the...
102

