Related Experiment Videos
Efficiency and Temporal Reasoning in Transformer-Based Video Object Detection: A Systematic Review
Yiannis Keravnos1,2, Anastasia Ioannou2, Andreas Papadopoulos1
1Cyprus Research and Innovation Center Ltd. (CYRIC), 2643 Nicosia, Cyprus.
Abstract:
This PRISMA 2020 systematic literature review analyzes efficient transformer-based video object detection (VOD). Across five databases through August 2025, 1343 records were identified, with 16 studies meeting the final inclusion criteria after full-text screening. Backward citation tracking added five additional eligible studies, yielding a small final evidence base of 21 studies (22 reports). This review's major contributions are (i) a focused narrative synthesis at the intersection of efficient model design, temporal modeling, and video object detection transformers, due to methodological heterogeneity in the corpus; (ii) ROB-CVA, a proposed risk-of-bias framework for computer vision; and (iii) an evidence-based analysis of temporal modeling strategies and evaluation practices. Furthermore, we examine evaluation inconsistencies, including dataset fragmentation and risk-of-bias. Fifty-two percent of studies were rated as low risk-of-bias using the review-specific, non-validated ROB-CVA framework, driven mainly by non-public code and datasets; this distribution is sensitive to our risk-of-bias aggregation rule. The certainty of evidence is low to moderate. The review outlines open challenges and future directions. This review is registered on OSF (doi: 10.17605/OSF.IO/Q9PJ4), supported by project AEOLUS (PHD IN INDUSTRY/1123/0145), and funded by the Republic of Cyprus through the Research and Innovation Foundation.