Related Experiment Video
Updated: Aug 5, 2026

Artificial Intelligence-Based System for Detecting Attention Levels in Students
Published on: December 15, 2023
Skeleton-Based Activity Recognition for Children with Autism Using Graph Convolutional Networks
Betül Ay1, Mehmet Ata Öztürk2, Galip Aydın1
1Department of Computer Engineering, Faculty of Engineering, Fırat University, 23119 Elazığ, Türkiye.
Abstract:
Movement-based and physical activity programs are central tools in autism intervention, so recognizing the activities a child performs during therapy is valuable for objective progress tracking. Manual monitoring of these sessions is time-consuming and subjective, and raw videos raise privacy concerns because it shows identifiable children. We address autism therapeutic activity recognition from privacy-preserving 2D skeletons, and we focus on the practical difficulty of how several therapeutic activities differ only in subtle motion details. As a backbone, we adopt ProtoGCN, a graph convolutional network that represents each action as a combination of learnable motion prototypes. However, this contrastive backbone organizes all classes at once, so it does not enforce a margin between the few pairs that remain entangled after training. We therefore introduce a Refine-Confusable (RC) module, a training-only regularizer that pushes apart the empirically most-confused class pairs using a hinge-margin loss over momentum-updated class centroids. The module changes neither the backbone nor the inference cost. On the MMASD dataset, restricted to the ten-class 2D-skeleton configuration, the RC module improves the base model across random, session-independent, and subject-independent evaluation. The gain is largest on the strictest subject-independent split and a clip-level analysis confirms that this improvement is statistically significant. Under the protocol-matched holdout, the method reaches 96.30% accuracy with 0.959 macro-F1, surpassing recent 2D-skeleton baselines while keeping a lightweight and privacy-preserving modality. The improvements are modest, as expected on a small clinical dataset, and t-SNE and prototype visualizations show that the learned representation is discriminative and interpretable.
