Related Experiment Videos
Imputation benchmarking in clinical databases: constraint enforcement for bounded neuropsychological assessments
Moad Hani1, Saïd Mahmoudi1, Mohammed Benjelloun1
1University of Mons, Faculty of Engineering, Belgium.
Background:
Missing data compromise the secondary use of clinical databases, particularly when variables are bounded by clinically defined ranges or discrete scoring rules. Imputation methods may achieve acceptable reconstruction error while producing clinically invalid values. We evaluated a constraint-aware imputation framework and examined whether broad patterns of method performance persisted across two independently collected Parkinson's disease cohorts.
Methods:
The primary benchmark used BD-Lille, with N = 1,357 participants and 51 bounded neuropsychological variables, and evaluated ten methods under three missingness mechanisms, four rates, and five prespecified seeds. A matched external evaluation used a cross-sectional baseline sample from the Parkinson's Progression Markers Initiative (PPMI), with N = 619 participants and seven bounded clinical variables, following the same factorial structure. PPMI reconstruction was assessed using macro-averaged range-normalised mean absolute error (NMAE), rank uncertainty, and paired constrained-unconstrained comparisons. An exploratory PPMI analysis assessed recovery of a method-independent complete-data partition at K = 3.
Results:
MissForest remained the most accurate method in the validated BD-Lille benchmark. In PPMI, EM achieved the lowest NMAE (0.05224, 95% CI 0.05194-0.05252), followed by HyperImpute (0.05244, 0.05212-0.05274) and MissForest (0.05399, 0.05364-0.05433). The numerical ordering therefore changed across cohorts. Although these three methods occupied the first three PPMI ranks in every bootstrap sample, the prespecified 0.01-NMAE practical-similarity margin classified eight methods within a broad competitive group. DAE and GAIN remained outside it. Constraint enforcement eliminated all PPMI range and measurement-support violations, with paired NMAE changes between -0.000392 and 0.000021. Fixed-reference ARI means ranged from 0.417 to 0.636 across methods, and reconstruction and partition-recovery rankings were related but not identical.
Conclusions:
The external evaluation did not support replication of a universal fine-grained ranking. It supported a more limited, context-dependent interpretation in which several classical, iterative, and ensemble methods remained competitive while their order changed. A versioned constraint layer enforced valid clinical ranges and measurement supports without a practically meaningful loss of reconstruction accuracy. The downstream findings were exploratory and indicate that reconstruction error and partition recovery should be evaluated as distinct outcomes.