Related Experiment Videos
Prediction of Atrial Fibrillation Occurrence With Handheld Mobile Electrocardiogram: Deep Learning Model Development
Minje Park1, Hyun Jin Ahn1, Yeongyeon Na1
1VUNO Inc., Seoul, Republic of Korea.
Background:
Atrial fibrillation (AF) is a common arrhythmia associated with an increased risk of stroke and heart failure. To improve prevention, recent studies have used deep learning models to identify at-risk individuals early from normal sinus rhythm (NSR). However, studies using mobile electrocardiogram (mECG) in outpatient, real-world settings remain underexplored.
Objective:
The study aimed to develop and evaluate deep learning models using a real-world limb-lead mECG database to predict the short-term occurrence of AF from NSR recordings.
Methods:
mECG data were collected from real-world users of commercially available handheld mECG devices capable of capturing 6 limb leads. AF occurrence was defined as an AF event within a predefined time window (7, 14, or 31 d) from the date of the NSR recording. Transformer-based prediction models were developed for limb-lead and lead I input configurations using a multistage training approach with self-supervised pretraining and domain adaptation, drawing on both open, large-scale clinical 12-lead ECG and proprietary real-world mECG databases. The models were evaluated in an internal real-world cohort and explored in an external cohort as a proof of concept via time-to-event analysis.
Results:
Between March 2023 and November 2024, 386,519 mECGs were acquired from 8206 users. There were 18,949, 25,206, and 33,524 AF incidences within the 7-, 14-, and 31-day time windows. The models were pretrained with 787,257 12-lead ECGs and 202,689 mECGs, then fine-tuned to predict AF occurrence using 97,447 labeled mECGs. The limb-lead models achieved areas under the receiver operating characteristic curves (AUROCs) of 0.793, 0.785, and 0.787 for 7-, 14-, and 31-day predictions on the internal cohort, respectively, with a user-level AUROC of 0.702 for the 31-day prediction. These models significantly outperformed the lead I models (P<.001), supporting the value of multilead configurations. The multistage pretraining was essential, as single-source pretraining yielded lower AUROCs of 0.555 with mECGs only and 0.761 with 12-lead ECGs only for the 31-day prediction. In the subgroup analysis, AUROC values were consistent across age, PR interval, and corrected QT interval, but showed disparities (P<.001) by sex (0.713 in females vs 0.794 in males) and by QRS duration (0.583 in ≥120 ms vs 0.796 in <120 ms). In the external cohort (n=144), the 31-day model stratified all 5 new-onset AF events, showing significantly different survival functions between the positively and negatively predicted groups (P=.03); Cox proportional hazards regression yielded a hazard ratio of 1.49 (95% CI 1.06-2.09) per 0.1 increase in model output.
Conclusions:
Our findings elucidate the feasibility of deep learning-based AF risk prediction using single-NSR recordings from mobile devices, highlighting the potential for remote AF management in real-world populations. The model output may serve as a risk indicator to support opportunistic AF screening, prompting further clinical evaluation and informing decisions about more intensive monitoring.