Related Experiment Video
Updated: Jan 10, 2026

06:37
Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention
Published on: December 15, 2023
5.2K
A tiny inertial transformer for human activity recognition via multimodal knowledge distillation and explainable AI.
Ismail Lamaakal1, Chaymae Yahyati1, Yassine Maleh2
1Multidisciplinary Faculty of Nador, Mohammed Premier University, Oujda, Morocco.
Scientific Reports
|November 27, 2025
Summary
XTinyHAR offers a lightweight, transformer-based solution for human activity recognition (HAR) on edge devices. This interpretable model achieves high accuracy and fast inference, making real-time HAR more accessible.
Area of Science:
- Computer Science
- Artificial Intelligence
- Machine Learning
Background:
- Accurate and interpretable human activity recognition (HAR) is crucial for healthcare, fitness, and smart environments.
- Deploying HAR models on resource-constrained edge devices presents significant challenges due to computational limitations.
Purpose of the Study:
- To introduce XTinyHAR, a novel, lightweight, transformer-based unimodal framework for human activity recognition.
- To enable efficient and interpretable HAR on edge devices through cross-modal knowledge distillation.
Main Methods:
- Developed XTinyHAR, a unimodal transformer framework utilizing temporal positional embeddings and attention rollout.
- Employed cross-modal knowledge distillation from a multimodal teacher model for training.
- Evaluated performance on UTD-MHAD and MM-Fit datasets.
Main Results:
- Achieved high test accuracies of 98.71% (UTD-MHAD) and 98.55% (MM-Fit) with matching F1-scores.
- Demonstrated excellent interpretability with Cohen's Kappa above 0.98.
- Model size is 2.45 MB, with fast inference (3.1 ms CPU, 1.2 ms GPU) and low computational cost (11.3M FLOPs).
Conclusions:
- XTinyHAR provides a high-performance, interpretable, and deployable solution for real-time HAR on edge devices.
- Ablation studies validated the effectiveness of individual model components.
- Subject-wise evaluations confirmed strong generalization capabilities across different users.