Related Experiment Video
Updated: Jun 27, 2025

A Step-by-Step Implementation of DeepBehavior, Deep Learning Toolbox for Automated Behavior Analysis
Published on: February 6, 2020
DeTAL: Open-Vocabulary Temporal Action Localization With Decoupled Networks.
This study introduces DeTAL, a novel two-stage method for open-vocabulary temporal action localization (OV-TAL). DeTAL improves zero-shot video understanding by decoupling detection and classification, outperforming existing approaches.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Machine Learning
Background:
- Pre-trained visual-language (ViL) models show promise in video understanding but struggle with open-vocabulary temporal action localization (OV-TAL) due to robustness issues.
- Adaptation methods like fine-tuning can cause misalignment between visual features and text descriptions for unseen actions, creating a detection-classification trade-off.
Purpose of the Study:
- To address the limitations of current ViL models in OV-TAL.
- To propose a robust and effective method that decouples action detection from classification for improved OV-TAL performance.
- To enable OV-TAL even without action category annotations during training.
Main Methods:
- Introduced DeTAL, a two-stage approach for OV-TAL.
- Decoupled action detection and action classification to avoid performance compromises.
- Adapted state-of-the-art close-set action localization methods for the OV-TAL task.
- Developed a new cross-dataset evaluation setting to assess zero-shot capabilities.
Main Results:
- DeTAL significantly improves performance in OV-TAL by avoiding the detection-classification trade-off.
- The method demonstrates effectiveness even when action category annotations are unavailable during training.
- Experimental results show DeTAL outperforms existing state-of-the-art methods on THUMOS14 and ActivityNet1.3 datasets.
Conclusions:
- DeTAL offers a simple yet effective solution for OV-TAL, enhancing robustness and performance.
- The proposed decoupling strategy successfully tackles the challenges of open-vocabulary action localization.
- DeTAL provides a strong baseline for future research in zero-shot video understanding and action localization.
More Related Videos
03:31Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
Published on: December 15, 2023
11:14A Novel Experimental and Analytical Approach to the Multimodal Neural Decoding of Intent During Social Interaction in Freely-behaving Human Infants
Published on: October 4, 2015
Related Concept Videos
Propagation of Action Potentials
Neurons (nerve cells) have a resting membrane potential, with a slightly negative charge inside compared to outside. This is maintained by ion channels, such as sodium (Na+) and potassium (K+) channels, which control the flow of ions. When a stimulus, like a touch or a signal from another neuron, triggers the neuron, sodium channels open, allowing sodium ions to...
State Space Representation
Consider an RLC circuit, a...
Multi-input and Multi-variable systems
In the absence...
State Space to Transfer Function
The transformation process begins with the state-space representation, characterized by the state equation and the output equation. These equations are typically represented as:
Associative Learning
Classical conditioning, also known...