Related Experiment Video
Updated: Sep 3, 2026

Real-World M3-BREATHE: Toward Multimodal Mobile Monitoring of Behaviour, Respiration, and Exposures for Treatment and Health Evaluation
Published on: June 5, 2026
Baseline normalization choices inflate classification performance in wearable health monitoring: quantification and
Alessandro Tognotti1,2, Corin F Otesteanu2, Luca Anceschi2
1Department of Electronics, Information and Bioengineering, Politecnico di Milano, Milan, Italy.
Abstract:
Wearable health monitoring systems provide continuous physiological measurements with growing use in clinical and population health applications. Baseline normalization enables individual-specific calibration of physiological signals to account for inter-subject variability; however, systematic evaluation of how the temporal relationship between normalization windows and evaluation data influences reported classification performance remains limited. Using physiological anxiety detection as a case study, we investigate this methodological challenge with the PhysioNet Spider Fear dataset (52 participants, including measurements of heart activity, skin conductance, and respiration), comparing normalization approaches and three baseline selection strategies. While maintaining strict subject-independent partitions, we show that including the normalized baseline in the evaluation window (test-inclusive baseline normalization) produced accuracies of 89%-93%, compared with when the baseline window is excluded (test-exclusive baseline normalization), which yielded 77%-85%, a 3-13 percentage-point difference consistent across classifiers, window sizes, and class-balance conditions. Two test-exclusive strategies using physiologically distinct baseline states (rest vs. anxiety) performed equivalently, indicating that performance differences were driven by temporal overlap rather than baseline selection. Full-feature normalization improved balanced accuracy by 11%-18% relative to unnormalized features; however, after removing temporal overlap, performance did not differ between normalization approaches. Prolonged stimulus exposure was associated with a 4-7 percentage-point reduction in balanced accuracy, consistent with habituation effects. Increasing window duration from 10 s to 60 s was associated with higher balanced accuracy (72%-74% vs. 78%-79%), whereas data augmentation did not improve performance under moderate class imbalance. We recommend explicit reporting of baseline temporal boundaries to enable accurate performance comparisons across studies and reliable assessment of clinical applicability in wearable health monitoring.