Related Experiment Video
Updated: Mar 15, 2026

08:25
Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
9.7K
Human Activity Recognition in Domestic Settings Based on Optical Techniques and Ensemble Models
Muhammad Amjad Raza1, Nasir Mehmood1, Hafeez Ur Rehman Siddiqui1
1Institute of Computer Science, Khwaja Fareed University of Engineering and Information Technology, Abu Dhabi Road, Rahim Yar Khan 64200, Punjab, Pakistan.
Sensors (Basel, Switzerland)
|March 14, 2026
Summary
This study introduces a vision-based human activity recognition (HAR) system using pose estimation and deep learning. The CNN-LSTM model achieved 98.78% accuracy, offering a precise and non-intrusive HAR solution.
Area of Science:
- Computer Vision
- Machine Learning
- Human-Computer Interaction
Background:
- Human Activity Recognition (HAR) is crucial for smart homes, healthcare, and assisted living.
- Traditional HAR methods using wearable sensors have limitations like user inconvenience and privacy concerns.
- Vision-based HAR offers a non-intrusive alternative but faces challenges with lighting and privacy.
Purpose of the Study:
- To develop and evaluate a novel, privacy-preserving, vision-based HAR system.
- To compare the performance of various deep learning architectures for HAR using skeletal keypoint data.
- To assess the robustness and generalization capabilities of the proposed HAR approach.
Main Methods:
- Utilized PoseNet for skeletal keypoint extraction from video data, creating privacy-preserving representations.
- Employed an ensemble of six deep learning models: Transformer, LSTM, GRU, MLP, 1D CNN, and CNN-LSTM.
- Evaluated models on 2734 activity samples from 30 subjects performing nine daily domestic activities.
Main Results:
- The Convolutional Neural Network-Long Short-Term Memory (CNN-LSTM) architecture achieved the highest accuracy of 98.78% on the test set.
- Leave-One-Subject-Out cross-validation demonstrated robust generalization, with CNN-LSTM yielding a mean accuracy of 97.21% ± 1.84%.
- Pose estimation effectively captured motion dynamics while preserving privacy, outperforming appearance-based methods.
Conclusions:
- Vision-based pose estimation combined with deep learning, particularly CNN-LSTM, provides a highly accurate and non-intrusive HAR solution.
- The proposed method is suitable for smart healthcare and home automation systems, addressing limitations of conventional approaches.
- The system demonstrates strong generalization capabilities, making it reliable for recognizing activities of unseen individuals.

