Related Experiment Video
Updated: Jun 24, 2026

Assessing the Accuracy of Fitness Smartwatch Data for Cardiovascular and Physical Activity Monitoring: A Validation Study in Digital Health
Published on: February 21, 2025
What matters beyond model choice for wearable sleep staging? How personalization, evaluation choices, and
Eric Canton1, Franco Tavella1, Christopher Drake2
1Arcascope Inc., Arlington, VA, United States.
Study Objectives:
Methods to improve sleep classification of wearable data often emphasize model choice. However, without widely used benchmark datasets, it is difficult to determine factors driving performance gains (model architecture, dataset selection, or evaluation decisions). This study examines how dataset characteristics influence sleep classification performance beyond model choice.
Methods:
We collected SleepAccel-Clinical, a dataset of Apple Watch acceleration and polysomnography from 28 individuals with sleep apnea. Neural network models were trained on this dataset alongside SleepAccel (31 healthy individuals) and DREAMT (100 individuals with suspected sleep disorders) to evaluate how training and testing set choices impact performance. All models were compared against an Easy To Classify (ETC) wake model, which estimates wake probability by smoothing and scaling activity in a brief time window. The ETC model's average area under the receiver operating characteristic curve (AUROC) on each dataset was used to quantify intrinsic ease of classification.
Results:
The ETC model accounted for substantial amount of the models' performance across the architectures, datasets, and training conditions, with significant correlations (p≪ .01) between ETC performance and all other scenarios. Including individuals with obstructive sleep apnea (OSA) in the training data improved performance when testing on datasets in which sleep disorders were suspected.
Conclusions:
Sleep classification performance depends heavily on training and testing dataset characteristics, not solely on model choice. The presence of ETC wake epochs can markedly inflate performance metrics. Researchers should benchmark models on widely available datasets and quantify intrinsic class separability when introducing new datasets. This article is part of the Consumer Sleep Technology Special Collection.
