Related Experiment Video
Updated: Jun 26, 2026

06:49
Evaluation of a Smartphone-based Human Activity Recognition System in a Daily Living Environment
Published on: December 11, 2015
9.4K
Deep Learning for Freezing of Gait Detection: Cross-Dataset Validation Reveals Critical Deployment Gaps Between
Wei Lin1,2, Sanjeet S Grewal2
1Department of Neurosurgery, The 904th Hospital of the Joint Logistics Support Force of People's Liberation Army, Wuxi 214000, China.
Sensors (Basel, Switzerland)
|February 27, 2026
Summary
Automated freezing of gait detection algorithms validated in labs fail in real-world Parkinson's disease patient use. A significant performance gap highlights the need for better training strategies for clinical translation.
Area of Science:
- Biomedical Engineering
- Neurology
- Machine Learning
Background:
- Freezing of gait (FoG) significantly impacts advanced Parkinson's disease (PD) patients, affecting 38-65%.
- Current automated FoG detection algorithms are primarily validated on limited laboratory datasets, raising concerns about their real-world clinical utility.
- A critical performance gap exists between laboratory-derived algorithm performance and its effectiveness in real-world, ecologically valid settings.
Purpose of the Study:
- To quantify the performance difference of automated FoG detection models trained and tested across laboratory and daily living datasets.
- To evaluate the impact of different training configurations, including class imbalance correction strategies, on model performance.
- To establish an empirical framework for assessing the deployment readiness of wearable FoG detection systems.
Main Methods:
- Temporal Convolutional Networks (TCNs) were employed to train FoG detection models.
- Models were trained on two distinct datasets: a daily living dataset (Figshare) and a laboratory dataset (DAPHNET).
- Five training configurations were compared, including different approaches to handle class imbalance (e.g., focal loss, weighting, sampling) and early stopping criteria (F1-score vs. Area Under the Curve).
Main Results:
- F1-based early stopping significantly outperformed Area Under the Curve (AUC)-based stopping (F1: 0.55 vs. 0.37).
- Combining multiple imbalance correction techniques paradoxically reduced precision due to excessive minority class over-weighting.
- A substantial 83% performance gap was observed between laboratory (F1: 0.9999 ± 0.0002) and daily living (F1: 0.55 ± 0.26) datasets, with a 1299-fold increase in variance.
Conclusions:
- Laboratory validation performance does not reliably predict real-world effectiveness for wearable FoG detection systems, indicating a "deployment gap".
- The deployment gap is influenced by environmental complexity, sensor limitations, and physiological variability inherent in daily living.
- Recommendations for training strategies are provided to improve the clinical translation and deployment readiness of FoG detection technologies.

