Related Experiment Video
Updated: Nov 17, 2025

07:09
Gaze in Action: Head-mounted Eye Tracking of Children's Dynamic Visual Attention During Naturalistic Behavior
Published on: November 14, 2018
11.2K
Learning to Recognize Actions on Objects in Egocentric Video With Attention Dictionaries
IEEE Transactions on Pattern Analysis and Machine Intelligence
|February 11, 2021
Summary
EgoACO, a novel deep neural architecture, enhances egocentric video action recognition by learning action-context-object descriptors. This approach achieves state-of-the-art performance by decoding key information from video frames.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Machine Learning
Background:
- Egocentric video action recognition is crucial for understanding human behavior.
- Existing methods often struggle to effectively integrate object and scene context with actions.
- Leveraging the verb-noun structure of action labels in egocentric datasets presents a unique challenge.
Purpose of the Study:
- To introduce EgoACO, a deep neural architecture for improved egocentric video action recognition.
- To develop a method that effectively pools action-context-object descriptors from frame-level features.
- To achieve state-of-the-art performance on benchmark egocentric action recognition datasets.
Main Methods:
- EgoACO utilizes a novel class activation pooling (CAP) layer, combining bilinear pooling and feature learning with self-attention.
- A recurrent version, Long Short-Term Attention (LSTA), is designed for temporal modeling, extending gated LSTMs with spatial attention.
- A multi-head prediction fuses action, object, and context descriptors, considering label inter-dependencies.
Main Results:
- EgoACO demonstrates state-of-the-art performance on the EPIC-KITCHENS and EGTEA Gaze+ datasets.
- The model successfully decodes action-context-object descriptors, significantly improving recognition accuracy.
- Built-in visual explanations aid in understanding the model's decision-making process.
Conclusions:
- EgoACO provides a powerful framework for egocentric video action recognition.
- The proposed CAP and LSTA mechanisms effectively capture essential action-context-object information.
- This work advances the field by achieving superior performance and offering interpretable insights.

