Related Experiment Video
Updated: Apr 12, 2026

Decoding Natural Behavior from Neuroethological Embedding
Published on: October 3, 2025
A multi-context fusion-aware graph modelling for group activity recognition using pose-conditioned spatial encoding
M R Tejonidhi1,2, K R Raghunandan3, B Uma4
1Nitte (Deemed to be University), NMAM Institute of Technology (NMAMIT), Department of Computer Science and Engineering, Nitte, 574110, Karnataka, India. tejonidhi.22phdecs214@student.nitte.edu.in.
Abstract:
Group activity recognition requires a holistic understanding of individual actions, their spatial relationships, and the surrounding environment. Traditional methods that focus solely on isolated movements often fail to capture the complex inter-player and scene-level dependencies inherent in sports and crowd scenarios. In this research work, a model for group activity recognition is developed. The proposed model combines various contextual features through the integration of poses of individual actors in the scene with the pose-aligned spatial scene context for relational reasoning. Pose features of individual actors are extracted using mmPose, while the scene-level context is encoded through pose-conditioned spatial feature aggregation rather than explicit semantic segmentation. These pose and scene context features extracted are combined and used to construct Actor Relation Graphs (ARGs) using Zero Normalized Cross Correlation (ZNCC) which improves robustness to appearance and variations in illumination. Further, Graph Convolutional Networks (GCNs) are modelled using relationships between individual actors in a scene and their group activities. The proposed framework explicitly combines pose-level and scene-level contextual features into a single relational graph, in contrast to previous ARG-GCN approaches that mainly rely on appearance features. The model is evaluated on two benchmark datasets: the Collective Activity dataset (CAD) and the Volleyball dataset (VD). The model exhibits classification accuracies of 95.02% and 94.81% on CAD and VD, respectively. On a TITAN-XP GPU, the average time per video clip with 41 frames is approximately 0.2 s. The results show that the combination of pose and scene contexts features enhances graph-based relational learning and improves recognition accuracy.
Related Concept Videos
Structural Classification of Joints
A fibrous joint is where the adjacent bones are united by fibrous connective...
Actor-Observer Effect
Functional Classification of Joints
The functional classification of joints is determined by the amount of mobility between the adjacent bones. Joints are functionally classified as a synarthrosis or immobile joint, an amphiarthrosis or slightly moveable joint, or as a diarthrosis, a freely moveable joint. Fibrous and cartilaginous joints can be functionally classified as either synarthroses or amphiarthroses, whereas all synovial joints are classified as diarthroses.
Synarthrosis
An...
Support Reactions in Three Dimensions
Ball and Socket Joint is one of the supports allowing free rotation about any axis. This freedom of rotation is...
Tagging and Fusion Proteins
Association Areas of the Cortex
Prefrontal Association Area: This area is located in the frontal lobe and is involved in planning, decision-making, and moderating social behavior. It connects with primary motor areas,...
