Related Experiment Video
Updated: Aug 5, 2026

03:31
End-To-End Deep Neural Network for Salient Object Detection in Complex Environments
Published on: December 15, 2023
Efficient Object Detection in Compressed Domain by Exploiting Knowledge Distillation from Pixel Domain
Serhat Dikyar1,2,3, Behcet Ugur Toreyin2,4
1ASELSAN Inc., Ankara 06370, Türkiye.
Journal of Imaging
|July 27, 2026
Summary
This study introduces a novel dual-phase framework for efficient object detection in compressed video data. It significantly reduces decoding latency and improves accuracy by processing data in the compressed domain, bypassing traditional bottlenecks.
Area of Science:
- Computer Vision
- Machine Learning
- Video Processing
Background:
- High-definition video data requires efficient real-time edge analytics.
- Traditional object detection relies on pixel-domain inputs, creating latency bottlenecks due to decoding.
- Compressed-domain processing offers a potential solution for faster analytics.
Purpose of the Study:
- To develop a fast and efficient object detection framework operating directly on partially decoded compressed-domain video data.
- To overcome the latency bottleneck associated with traditional pixel-domain object detection pipelines.
- To improve the efficiency and accuracy of object detection for edge analytics.
Main Methods:
- Proposed a dual-phase framework for compressed-domain object detection.
- Introduced Low-Frequency Spectral Prioritization for partial decoding, reducing data and accelerating processing.
- Employed multi-granularity cross-domain knowledge distillation to transfer knowledge from a pixel-domain teacher to a compressed-domain student network.
Main Results:
- The proposed framework enables object detection directly on partially decoded HEVC (High Efficiency Video Coding) data.
- Experiments on COCO-mini dataset with RetinaNet, FCOS, and GFL showed improved mAP (+0.96%) compared to pixel-domain baselines.
- Achieved reduced average decoding latency by retaining fundamental residual coefficients.
Conclusions:
- The novel dual-phase framework offers a viable solution for efficient, low-latency object detection in compressed video streams.
- Compressed-domain processing, combined with knowledge distillation, effectively mitigates traditional decoding bottlenecks.
- This approach enhances the feasibility of real-time edge analytics for high-definition video data.