Related Experiment Videos
Systematic benchmark of reduced-lead configurations for 12-lead ECG reconstruction: multi-model evaluation across all
Xinyu Zhang1, Hailing Cai1, Huilong Duan1
1College of Biomedical Engineering and Instrument Science, Zhejiang University, Hangzhou, Zhejiang, China.
Abstract:
Portable reduced-lead ECG devices offer a practical path to pre-hospital emergency care and community cardiovascular screening, yet a foundational question has remained unresolved: which subset of the standard 12 leads should such a device actually record? Prior work evaluates a small number of fixed configurations selected by convention, without systematic comparison across the full combinatorial space or adequate accounting for the true cost of electrode attachment. We present the first exhaustive benchmark evaluating all 4,094 C(12, N) lead subsets for N = 1-11 under four reconstruction paradigms-linear regression, ridge regression, a lightweight 1-D convolutional network, and a Transformer encoder-decoder-on the PTB-XL dataset. Performance is assessed along three orthogonal axes: reconstruction fidelity (PCC, RMSE, SNR on withheld leads), downstream diagnostic accuracy (macro-F1 across five cardiac superclasses via a classifier trained on real 12-lead signals), and acquisition efficiency operationalised as electrode contact burden. A Composite Lead Score with five-point α-sensitivity analysis identifies N = 4 as the consensus efficiency-accuracy knee (mean macro-F1 = 0.631, 93.5% of the 12-lead upper bound). External zero-shot validation on CPSC2018 (6,877 records) and Chapman-Shaoxing (45,152 records) confirms that all three deployment-recommended configurations retain ≥ 83% of within-PTB-XL three-class F1 on both cohorts (Chapman 92-99%; CPSC 83-91%), and the N = 3 knee replicates across all four architectures on Chapman, supporting interpretation of the knee as a property of the lead system's information geometry rather than of any specific architecture or development cohort. A frozen-classifier sensitivity control (ΔF1 = + 0.0006 at N = 4 after fine-tuning, an order of magnitude smaller than the cross-N marginal gain) confirms the benchmark's robustness to downstream-model assumptions. Evidence-based configurations are derived for three deployment scenarios: V6 for pre-hospital triage, I + II + AVR + AVF for community screening, and a 7-lead set at 5 contacts for home monitoring; we present these as population-level defaults whose within-tier specific lead choice may benefit from cohort-specific re-derivation. All code and benchmark results are publicly released.