Related Experiment Video
Updated: May 2, 2026

A Step-by-Step Implementation of DeepBehavior, Deep Learning Toolbox for Automated Behavior Analysis
Published on: February 6, 2020
LSGNet: A Local-Pattern Separation and Global-Aware Network for Temporal Action Detection
Abstract:
Temporal Action Detection aims to localize and classify action instances within untrimmed videos, yet it remains challenging due to background clutter, high intra-class similarity, and varied temporal scales in real-world scenarios. To address these issues, we propose the Local-Pattern Separation and Global-Aware Network (LSGNet) tailored for temporal action localization. Specifically, the core of LSGNet is the Local Pattern Separation Module (LPSM), which explicitly models both consistency and variation patterns of action segments within local temporal windows. Additionally, to capture comprehensive contextual information, we introduce the Global Context-Aware Representation Module (GCRM), which decouples temporal features across multiple granularities and enables robust modeling of long-range dependencies. Finally, we design the Multi-scale Feature Refinement Module (MFRM) to mitigate the degradation of fine-grained information by performing iterative reconstruction across temporal scales, thereby enriching semantic representations and preserving temporal details. Extensive experiments on THUMOS14, ActivityNet1.3, HACS, and EPIC-Kitchens-100 demonstrate the effectiveness of the proposed LSGNet method. Additional ablation studies on the QVHighlights dataset further confirm the generalization capability of LPSM module in video moment retrieval and highlight detection, achieving consistent improvements in retrieval accuracy and localization precision.