Related Experiment Video
Updated: Oct 2, 2026

Real-World M3-BREATHE: Toward Multimodal Mobile Monitoring of Behaviour, Respiration, and Exposures for Treatment and Health Evaluation
Published on: June 5, 2026
Incremental Value of Smartphone Sensing for Monitoring Momentary Affect Intensity in Adults Using Transformer-Based
Yiqin Zhu1, Yuyi Yang2, Renee J Thompson1
1Department of Psychological and Brain Sciences, Washington University in St. Louis, 1 Brookings Drive, CB 1125, St. Louis, MO, 63130, United States, +1-314-935-3502.
Background:
Ubiquitous smartphone access and statistical advances offer opportunities to continuously track affect intensity, which is central to various psychological processes and behaviors. Research demonstrated the potential of personalized predictions of momentary negative affect (NA) and positive affect (PA) using passive sensing. However, studies typically incorporated all available data sources without differentiating their added value, nor did they investigate whether refining location features with self-reported semantic location (eg, workplaces) improved personalized predictions.
Objective:
We evaluated three specific aims: (1) how different combinations of data sources improved performance compared to personalized baseline models, (2) whether model predictions differed across passive data aggregation timescales, and (3) whether incorporating self-reported semantic locations improved model predictions.
Methods:
Adults (final n=133) completed a 14-day ecological momentary assessment (EMA) protocol reporting emotional experiences 5 times daily alongside smartphone sensing. Testing data (n=532 EMAs) used the last 4 surveys for each individual, with the remaining used for training and validation (n=6805 [NA]/6800 [PA] EMAs). We evaluated whether combinations of personalization, passive sensing, and affect history improved baseline prediction, and how full-information temporal fusion transformers (TFTs) performed across 6 timescales (1, 3, 6, 12, 24, and 48 hours), with or without self-reported semantic location features.
Results:
The baseline model, using each individual's mean affect in the training set, demonstrated moderate predictive performance for NA (mean absolute error [MAE]=0.66, 95% CI 0.60-0.73; R²=40.2%) and PA (MAE=0.71, 95% CI 0.65-0.78; R²=36.1%). Full-information TFTs improved NA prediction (MAE=0.62, 95% CI 0.56-0.69; R²=45.2%; ΔMAE=-0.04, 95% CI -0.06 to -0.01; P values ≤.004; Cohen d=-0.27) but not PA prediction (MAE=0.70, 95% CI 0.64-0.78; R²=32.5%; ΔMAE=-0.01, 95% CI -0.03 to 0.02; P values >.10; Cohen d=-0.04). No pairwise timescale comparison survived false discovery rate (FDR) correction (NA: PFDR=.05-.98; PA: PFDR=.08-.99). Adding self-reported locations did not improve NA prediction (ΔMAE=0.02, 95% CI -0.01 to 0.04; P values >.20; Cohen d=0.11) or PA prediction (ΔMAE=0.01, 95% CI -0.01 to 0.03; P values>.33; Cohen d=0.07). However, incorporating self-reported semantic locations changed the composition and relative ranking of important inputs, with these changes varying across NA and PA and between past and future inputs.
Conclusions:
Incorporating smartphone features provided a modest and significant improvement in momentary NA prediction, but not PA prediction. Model performance did not vary across passive data aggregation timescales. While adding self-report semantic locations did not improve prediction accuracy, it changed variable-importance patterns and may provide additional context for interpreting digital behavioral markers. Future personalized predictions should incorporate person-mean affect as an essential benchmark. These findings support passive smartphone sensing as a valuable supplement to, rather than a replacement for, active EMA.
