Related Experiment Video
Updated: Sep 28, 2026

Methodology for Establishing a Community-Wide Life Laboratory for Capturing Unobtrusive and Continuous Remote Activity and Health Data
Published on: July 27, 2018
Duplicate leakage and evaluation integrity on a public synthetic sleep-health tabular dataset: A methodological case
Pramod K B Rangaiah1, B P Pradeep Kumar2, Robin Augustine1
1Microwaves in Medical Engineering Group, Division of Integrated Smart Systems Technology, Department of Electrical Engineering, Uppsala University, Box 65, SE-751 03, Uppsala, Sweden.
Abstract:
Public tabular datasets are widely reused to benchmark machine learning for health-related screening, but their evaluation validity is rarely audited. Using the widely cited, synthetic Sleep Health and Lifestyle Dataset (374 records, three classes) as a case study, we show that the very high accuracies reported for it reflect evaluation optimism rather than a property of the prediction task: several published studies report accuracies approaching 98%, whereas our reproduction under conventional cross-validation reaches approximately 91%, and confining duplicate records to a single fold reduces this further. Our central finding is duplicate-record leakage: only 132 of the 374 rows are distinct, 64.7% of rows are exact copies of another row, and the conventional cross-validation used in prior work allows identical rows to fall on both sides of the train/test split. When duplicate records are confined to a single fold using group-aware cross-validation, a gradient-boosted model drops from 91.2% to 72.6%, and on fully de-duplicated data to 61.7%: standard cross-validation overestimates performance by 18.6 percentage points relative to duplicate-group-aware evaluation, and performance falls by a further 10.9 percentage points after full de-duplication, reflecting the reduced effective sample size and conflicting-label structure. Under the corrected protocol we benchmark seven classifiers with correlation-corrected (Nadeau-Bengio) confidence intervals, nested cross-validation for hyperparameter selection, and effect-size-oriented significance testing. Simple models (logistic regression, support vector machine, multilayer perceptron) and CatBoost remain robust (≈87-91%), whereas the untuned random forest, XGBoost, and LightGBM show substantial variability, with very wide intervals; after nested tuning they partially recover, underscoring that model rankings on this dataset are not reliable. As a secondary, deliberately negative result, a feature-graph attention network we designed for the task (with a learnable sparse graph and a residual gate) does not outperform a plain multilayer perceptron (full model 88.2%±4.2 vs 89.3%±3.4; paired Wilcoxon p=0.03), indicating that graph inductive biases do not provide measurable benefit for this small static tabular dataset. We also report a near-deterministic association between the Occupation field and the label (Cramér's V=0.75; an 85% no-learning lookup), which we treat as suggestive of a data-construction artefact rather than as proven leakage. We conclude, in line with recent leakage-aware work in medical imaging, that this dataset is unsuitable for strong claims about sleep-disorder screening unless duplicate structure and synthetic-construction artefacts are handled explicitly; the contribution is a reusable, duplicate-aware benchmarking protocol rather than a clinical result.