Related Experiment Video
Updated: Aug 22, 2025

08:25
Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
9.1K
Attention-Guided Disentangled Feature Aggregation for Video Object Detection
Shishir Muralidhara1,2, Khurram Azeem Hashmi1,2,3, Alain Pagani3
1Department of Computer Science, Technical University of Kaiserslautern, 67663 Kaiserslautern, Germany.
Sensors (Basel, Switzerland)
|November 11, 2022
Summary
This study introduces an attention-heavy framework for video object detection, improving accuracy by disentangling and aggregating frame features. The novel approach enhances object localization and classification in challenging video sequences.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Machine Learning
Background:
- Object detection in still images is well-established, but video object detection faces challenges like blur and occlusion.
- Existing methods struggle with the dynamic nature of video data, impacting localization and classification accuracy.
Purpose of the Study:
- To propose an attention-heavy framework for robust video object detection.
- To address challenges in video object detection by effectively aggregating frame-level features.
- To improve the performance of object detection in video sequences.
Main Methods:
- A two-stage object detection framework based on Faster R-CNN architecture.
- Integration of scale, spatial, and task-aware attention in a disentanglement head.
- Utilization of temporal attention in an aggregation head to combine features from support frames.
Main Results:
- Achieved a mean Average Precision (mAP) of 49.8 with ResNet-50 backbone.
- Achieved a mean Average Precision (mAP) of 52.5 with ResNet-101 backbone on the ImageNet VID dataset.
- Demonstrated significant performance improvement over individual baseline methods.
Conclusions:
- The proposed attention-heavy framework effectively tackles challenges in video object detection.
- Feature disentanglement and temporal aggregation significantly enhance detection accuracy.
- The approach offers a promising solution for real-world video analysis applications.
Related Concept Videos
Aggregates Classification
366
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
366
Association Areas of the Cortex
5.8K
Association areas are regions of the cerebral cortex that do not have a specific sensory or motor function. Instead, they integrate and interpret information from various sources to enable higher cognitive processes such as memory, learning, and decision-making. Some key association areas include the following:
Prefrontal Association Area: This area is located in the frontal lobe and is involved in planning, decision-making, and moderating social behavior. It connects with primary motor areas,...
Prefrontal Association Area: This area is located in the frontal lobe and is involved in planning, decision-making, and moderating social behavior. It connects with primary motor areas,...
5.8K

