Related Experiment Video
Updated: Sep 16, 2025

04:23
A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
Published on: April 21, 2023
2.0K
A Review of DEtection TRansformer: From Basic Architecture to Advanced Developments and Visual Perception
Liang Yu1, Lin Tang1, Lisha Mu1
1College of Software Engineering, Sichuan Polytechnic University, Deyang 618000, China.
Sensors (Basel, Switzerland)
|July 12, 2025
Summary
This paper reviews the evolution of Detection Transformers (DETR) for object detection. It details improvements in attention, queries, training, and efficiency, addressing DETR
Area of Science:
- Computer Vision
- Artificial Intelligence
- Machine Learning
Background:
- Detection Transformer (DETR) pioneered end-to-end object detection using Transformers, removing manual components like anchor boxes and Non-Maximum Suppression (NMS).
- Original DETR faced challenges including slow convergence, inadequate small object detection, and low computational efficiency, necessitating further research and development.
Purpose of the Study:
- To systematically review the technical evolution of DETR models from a problem-driven perspective.
- To analyze advancements in core components such as attention mechanisms, query design, training strategies, and architectural efficiency.
- To provide a comprehensive overview of DETR's applications and future research directions.
Main Methods:
- A systematic literature review focusing on problem-driven advancements in DETR.
- Analysis of key technical improvements including attention mechanisms, query formulation, and training methodologies.
- Categorization of DETR applications across various domains like autonomous driving and medical imaging.
Main Results:
- Significant progress has been made in improving DETR's convergence speed, small object detection accuracy, and overall efficiency.
- Key advancements have been observed in attention mechanisms, query optimization, and training strategies.
- DETR has demonstrated broad applicability in diverse fields, including autonomous driving, medical imaging, and remote sensing.
Conclusions:
- The DETR architecture has evolved substantially, addressing its initial limitations through targeted innovations.
- Continued research is exploring extensions to fine-grained classification and video understanding.
- This review provides a structured understanding of DETR's development, highlighting current challenges and future research avenues.
Related Concept Videos
Visual System
695
Light enters the eye through the cornea, a transparent, dome-shaped surface covering the surface of the eyeball that helps to direct and focus incoming light. This light is then channeled toward the pupil, an adjustable opening whose size is controlled by the iris. The iris, a pigmented muscle, regulates the amount of light entering the eye by contracting or dilating the pupil, thereby ensuring optimal light levels for clear vision.
Once through the pupil, the light passes through the lens, a...
Once through the pupil, the light passes through the lens, a...
695
Vision
55.4K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
55.4K
Parallel Processing
234
The brain processes sensory information rapidly due to parallel processing, which involves sending data across multiple neural pathways at the same time. This method allows the brain to manage various sensory qualities, such as shapes, colors, movements, and locations, all concurrently. For instance, when observing a forest landscape, the brain simultaneously processes the movement of leaves, the shapes of trees, the depth between them, and the various shades of green. This enables a quick and...
234
Three-Winding Transformers
314
Three identical single-phase transformers can be configured to form a three-phase transformer connection, which involves high-voltage and low-voltage windings. The high-voltage windings are denoted by capital letters A-B-C, while the low-voltage windings are labeled with lowercase letters a-b-c, representing their respective phases. This notation helps distinguish between the high and low voltage sides of the transformer.
In the per-unit equivalent circuit of a grounded Y-Y three-phase...
In the per-unit equivalent circuit of a grounded Y-Y three-phase...
314
Types Of Transformers
1.1K
Transformers can provide desired voltages to a circuit by modifying the number of turns in the secondary windings.
If the ratio of the number of turns in the secondary winding to that of the primary winding is greater than one, then the transformer is said to be a step-up transformer. In a step-up transformer, the voltage at the secondary winding is greater than the voltage applied at the primary winding.
However, if this ratio is less than one, the transformer is said to be a step-down...
If the ratio of the number of turns in the secondary winding to that of the primary winding is greater than one, then the transformer is said to be a step-up transformer. In a step-up transformer, the voltage at the secondary winding is greater than the voltage applied at the primary winding.
However, if this ratio is less than one, the transformer is said to be a step-down...
1.1K
Depth Perception and Spatial Vision
944
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
944

