Related Experiment Video
Updated: May 26, 2026

End-To-End Deep Neural Network for Salient Object Detection in Complex Environments
Published on: December 15, 2023
A scalable multimodal framework for learning engagement recognition using three-dimensional convolutional neural
Kuan-Cheng Lin1, Chiung-Chen Tseng1, Junyi Wu1
1Department of Management Information Systems, National Chung Hsing University, Taichung, Taiwan.
Abstract:
This study developed a learning engagement recognition system that integrates behavioral and emotional modalities. By employing a three-dimensional (3D) convolutional neural network (CNN) model and a semi-automatic annotation mechanism, the developed system accurately analyzes students' nonverbal behaviors and facially displayed emotions during class to identify their level of engagement. The study employed 570 instructional video recordings, segmented into 44,059 clips of 10s each. Three experts annotated the levels of emotional and behavioral engagement shown in these clips on the basis of standardized guidelines. Two separate 3D CNN models were trained to recognize emotional and behavioral features. Under the revised video-level validation setting, the emotional engagement model achieved an overall accuracy of 0.92, whereas the behavioral engagement model achieved an overall accuracy of 0.80. Moreover, Fleiss' kappa was used to evaluate the consistency between model predictions and human annotations, indicating "almost perfect agreement." The results demonstrate that integrating 3D convolutional neural networks with standardized, semi-automatic annotation rules can substantially enhance both the accuracy and scalability of automated learning engagement recognition systems while maintaining human-level reliability. The proposed framework provides a robust foundation for scalable AI-driven learning analytics and related video-based behavior recognition applications.
