Related Experiment Video
Updated: Aug 15, 2026

Estimating Bilateral Atrial Function by Cardiovascular Magnetic Resonance Feature Tracking in Patients with Paroxysmal Atrial Fibrillation
Published on: July 20, 2022
RR-interval-based atrial fibrillation detection and burden estimation: cross-dataset validation and calibration-aware
Mohammad Haekal1, Siti Nurul Khotimah1, Galih Restu Fardian Suwandi1
1Nuclear Physics and Biophysics Research Group, Faculty of Mathematics and Natural Sciences, Institut Teknologi Bandung, Bandung, Indonesia.
Insights
Accurate atrial fibrillation (AF) burden estimation requires calibrated AF detection probabilities. Recalibration significantly improved probability accuracy and burden estimation in external datasets, highlighting its importance for reliable AF burden assessment.
Area of Science:
- Cardiology
- Biomedical Engineering
- Machine Learning in Healthcare
Background:
- Atrial fibrillation (AF) burden is a critical endpoint in long-term rhythm monitoring.
- Reliable AF burden estimation depends not only on accurate AF detection but also on probability calibration, especially under external dataset shifts.
- Existing methods may face challenges in maintaining burden validity when predicted AF probabilities are aggregated over time.
Purpose of the Study:
- To develop an interpretable RR-interval feature model for AF detection.
- To evaluate the model's performance, including probability calibration, on development and independent external datasets.
- To assess the impact of probability calibration on AF burden estimation derived from aggregated probabilities.
Main Methods:
- Developed an RR-interval feature model for AF detection.
- Evaluated the model using record-wise cross-validation and independent external validation on public Holter ECG databases.
- Assessed window-level performance using ROC-AUC, PR-AUC, Brier score, ECE, and calibration metrics; estimated recording-level AF burden using probability-based and hard-label aggregation.
Main Results:
- The model demonstrated high discrimination (external ROC-AUC: 0.9868, PR-AUC: 0.9872) but showed deteriorated external calibration (ECE(15): 0.1219).
- Probability-based AF burden estimation showed strong association with reference burden but weaker agreement than hard-label aggregation (MAE: 0.1473 vs. 0.0836), indicating systematic probability underprediction.
- External recalibration (Platt and isotonic) substantially improved probability quality (median ECE(15) reduced to 0.0394 and 0.0287) and probability-based burden estimation (median MAE reduced to 0.0692 and 0.0604).
Conclusions:
- RR-interval-based AF detection maintains strong ranking performance across datasets.
- Probability calibration is crucial for the validity of AF burden estimates derived from aggregated probabilities.
- Explicit evaluation and application of recalibration methods are recommended when using predicted probabilities for AF burden estimation.
Abstract:
Objective.Atrial fibrillation (AF) burden has become an increasingly important endpoint in long-duration rhythm monitoring, but reliable burden estimation requires more than accurate AF detection alone. In particular, when burden is derived by aggregating predicted AF probabilities over time, probability calibration may directly affect burden validity under external dataset shift.Approach.This study developed an interpretable-interval feature model for AF detection and evaluated it using record-wise cross-validation on a development cohort and independent cross-dataset external validation on public Holter electrocardiographic databases. Window-level performance was assessed using the area under the receiver operating characteristic curve (ROC-AUC), area under the precision-recall curve (PR-AUC), Brier score, expected calibration error (ECE), and calibration intercept and calibration slope. Recording-level AF burden was estimated using both probability-based and hard-label aggregation and evaluated using mean absolute error (MAE) and agreement analyses.Main results.The model showed high discrimination in both development and external evaluation, with external ROC-AUC ofand PR-AUC of. However, external calibration deteriorated despite preserved ranking performance, with Brier score of, ECE(15) of, calibration intercept of, and calibration slope of. In the external cohort, probability-based burden estimation preserved strong association with reference burden but showed weaker raw agreement than hard-label aggregation, with MAE ofversus, consistent with systematic probability underprediction. Repeated external recalibration across record-level splits substantially improved probability quality and probability-based burden estimation. Median probability-burden MAE decreased fromwithout recalibration toafter Platt recalibration andafter isotonic recalibration, while median ECE(15) decreased fromtoand, respectively.Significance.These findings indicate that-interval-based AF detection maintained strong ranking performance in the tested external cohort, but probability calibration should be evaluated explicitly when predicted probabilities are aggregated into AF-burden estimates.