Related Experiment Video
Updated: Aug 2, 2025

Development of a Gaze-Contingent Display Framework Designed for Perceptual and Oculomotor Research with Simulated Central Vision Loss
Published on: April 11, 2025
PLG-ViT: Vision Transformer with Parallel Local and Global Self-Attention
Nikolas Ebert1,2, Didier Stricker2, Oliver Wasenmüller1
1Research and Transfer Center CeMOS, Mannheim University of Applied Sciences, 68163 Mannheim, Germany.
This study introduces the Parallel Local-Global Vision Transformer (PLG-ViT), a novel architecture that efficiently combines local and global self-attention for computer vision tasks. PLG-ViT achieves superior performance in image classification and segmentation compared to existing models.
Area of Science:
- Computer Vision
- Deep Learning
- Artificial Intelligence
Background:
- Transformer architectures, particularly those utilizing self-attention, have surpassed Convolutional Neural Networks (CNNs) in various computer vision applications.
- Self-attention mechanisms allow transformers to capture dependencies across both short and long distances, creating extensive receptive fields.
Purpose of the Study:
- To propose the Parallel Local-Global Vision Transformer (PLG-ViT), a versatile backbone model designed to enhance computer vision performance.
- To effectively and efficiently represent short- and long-range spatial interactions by merging local and global self-attention features.
Main Methods:
- Developed the Parallel Local-Global Vision Transformer (PLG-ViT) by integrating local window self-attention with global self-attention.
- Evaluated PLG-ViT on image classification, object detection, instance segmentation, and semantic segmentation tasks.
- Compared PLG-ViT against CNN-based and state-of-the-art transformer-based architectures, including ConvNeXt and Swin Transformer.
Main Results:
- PLG-ViT demonstrated superior performance over CNN and existing transformer models in various computer vision tasks.
- Achieved high Top-1 accuracy on ImageNet-1K: 83.4% (27M parameters), 84.0% (52M parameters), and 84.5% (91M parameters).
- Outperformed similarly sized networks like ConvNeXt and Swin Transformer.
Conclusions:
- The proposed PLG-ViT effectively fuses local and global self-attention for efficient and powerful visual representation.
- PLG-ViT offers a competitive and efficient alternative to current leading computer vision backbones.
- The model shows significant potential for complex downstream tasks in computer vision.
Related Concept Videos
Parallel Processing
Types Of Transformers
If the ratio of the number of turns in the secondary winding to that of the primary winding is greater than one, then the transformer is said to be a step-up transformer. In a step-up transformer, the voltage at the secondary winding is greater than the voltage applied at the primary winding.
However, if this ratio is less than one, the transformer is said to be a step-down...
The Ideal Transformer
Ampere's Law forms the basis of understanding the magnetic field within the transformer. It states that the integral of the magnetic field intensity's...
Equivalent Circuits for Practical Transformers
In a practical transformer, each winding exhibits resistance and leakage reactance. The...
Three-Winding Transformers
In the per-unit equivalent circuit of a grounded Y-Y three-phase...
Transformers in Distribution System
Distribution substation transformers come in various ratings and typically use mineral oil for insulation and cooling. To prevent moisture and air from entering the oil, some transformers use an inert gas like nitrogen to fill the...

