Related Experiment Video
Updated: Jan 11, 2026

04:23
A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
Published on: April 21, 2023
2.3K
PMRVT: Parallel Attention Multilayer Perceptron Recurrent Vision Transformer for Object Detection with Event Cameras
Zishi Song1, Jianming Wang2, Yongxin Su3
1School of Electronics and Information, Tiangong University, Tianjin 300387, China.
Sensors (Basel, Switzerland)
|November 13, 2025
Summary
This study introduces PMRVT, a novel framework for event-based object detection. PMRVT enhances real-time performance and accuracy in dynamic environments by balancing efficiency, spatial expressiveness, and temporal consistency.
Area of Science:
- Computer Vision
- Robotics
- Artificial Intelligence
Background:
- Object detection in high-speed, dynamic environments is challenging for conventional cameras due to motion blur and latency.
- Event cameras offer an alternative with high dynamic range and low latency, but existing methods have high computational costs and limited temporal modeling.
- Real-time event-based vision applications require improved detection accuracy and efficiency.
Purpose of the Study:
- To present PMRVT (Parallel Attention Multilayer Perceptron Recurrent Vision Transformer), a unified framework for event-based object detection.
- To systematically balance early-stage efficiency, spatial expressiveness, and long-horizon temporal consistency in event-based detection.
- To achieve strong accuracy and real-time performance for dynamic environments.
Main Methods:
- Developed a hybrid hierarchical backbone for efficient processing.
- Implemented a Parallel Attention Feature Fusion (PAFF) mechanism with a coordinated dual-path design.
- Employed a temporal integration strategy for enhanced consistency.
Main Results:
- PMRVT achieved 48.7% mAP on the Gen1 dataset and 48.6% mAP on the 1 Mpx dataset.
- Inference latencies were 7.72 ms (Gen1) and 19.94 ms (1 Mpx).
- Demonstrated a 1.5 pp accuracy improvement and 8% latency reduction compared to state-of-the-art methods.
Conclusions:
- PMRVT offers a favorable balance between accuracy and speed for event-based object detection.
- The framework provides a reliable solution for real-time event-based vision applications.
- PMRVT addresses limitations of existing methods in computational cost and temporal modeling.
Related Concept Videos
Parallel Processing
612
The brain processes sensory information rapidly due to parallel processing, which involves sending data across multiple neural pathways at the same time. This method allows the brain to manage various sensory qualities, such as shapes, colors, movements, and locations, all concurrently. For instance, when observing a forest landscape, the brain simultaneously processes the movement of leaves, the shapes of trees, the depth between them, and the various shades of green. This enables a quick and...
612
Vision
59.3K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
59.3K
Visual System
1.6K
Light enters the eye through the cornea, a transparent, dome-shaped surface covering the surface of the eyeball that helps to direct and focus incoming light. This light is then channeled toward the pupil, an adjustable opening whose size is controlled by the iris. The iris, a pigmented muscle, regulates the amount of light entering the eye by contracting or dilating the pupil, thereby ensuring optimal light levels for clear vision.
Once through the pupil, the light passes through the lens, a...
Once through the pupil, the light passes through the lens, a...
1.6K
