Related Experiment Videos
Label-efficient cross-population transfer learning for electrocardiographic risk stratification: evaluation across
1Department of Functional Examination, The Fifth Affiliated Hospital of Anhui University of Chinese Medicine, Lu'an, Anhui, China.
Background:
Electrocardiography (ECG) is the most widely deployed cardiac screening modality, but demographic, acquisition and label-space heterogeneity across cohorts obstructs clinical translation. Cross-population studies rarely quantify which conditions transfer, how much target labelling is needed, or whether the adapted score stays calibrated and clinically useful.
Methods:
We harmonised 12,521 recordings from PTB-XL (Germany; 7,371 records, 6,643 patients), Chapman-Ningbo (China; 4,574) and the MIT-BIH Arrhythmia Database (USA; 576 segments, 48 recordings) into 13 SNOMED-CT super-classes. A 38-dimensional interpretable feature set, computed in physical units and standardised on the source training partition alone, fed a compact multi-task backbone (30,400 parameters) trained on patient-separated partitions. It was evaluated zero-shot on both targets and adapted with 5%-50% of Chapman-Ningbo labels by linear probing on the frozen embedding, with partial and full fine-tuning as comparators; prevalence-aware sampling made four target classes (NORM, STD, LVH, TINV) evaluable. The composite diagnostic score was recalibrated on a dedicated split and assessed on untouched test data using bootstrap intervals, Holm-corrected comparisons and subgroup interaction tests.
Results:
The source backbone reached macro-AUROC 0.866 (0.837-0.891) across 13 classes. Zero-shot transfer recovered 88% of the four-class source reference on Chapman-Ningbo (0.754, 0.727-0.779) and 0.642 (0.546-0.735, record-clustered) on MIT-BIH. Linear probing reached 0.795 with 5% of target labels and 0.829 with 50%, below the four-class source reference of 0.859; after Holm correction the gain over zero-shot was significant for NORM (+0.135), STD (+0.093) and TINV (+0.065) but not LVH (+0.004). Full fine-tuning added 0.021 at 50% labels but was inferior at 5% and less well calibrated. The recalibrated composite score attained AUROC 0.791 (0.743-0.836), Brier score 0.055 and expected calibration error 0.018 on Chapman-Ningbo, where 97% of positives were left-ventricular hypertrophy, and 0.834 on the seven-condition PTB-XL composite. Discrimination varied across age strata (0.708-0.863; interaction p = 0.012).
Conclusion:
Cross-population transfer of ECG models is feasible with label-efficient fine-tuning of a compact, interpretable feature backbone, and the resulting composite diagnostic score remains well-calibrated in a geographically distinct population. Subgroup heterogeneity nonetheless persists, with lower discrimination in participants aged 65 years and over and a sex-dependent calibration slope, and requires further validation.