Related Experiment Videos
Evaluating the Cyclical Hybrid Imputation Technique (CHIT) under MCAR, MAR, and MNAR missing data mechanisms:
Kurban Kotan1, Serdar Kırışoğlu2
1Department of Artificial Intelligence and Machine Learning, Faculty of Computer and Information Sciences, Manisa Celal Bayar University, Manisa, Turkey.
Abstract:
Choosing an appropriate imputation strategy for clinical datasets requires understanding which missing data mechanism is operative, yet most imputation benchmarks conflate performance across mechanisms. The present work addresses this gap by conducting the first mechanism-stratified comparison of three imputation approaches-CHIT, Multiple Imputation by Chained Equations via BayesianRidge (MICE), and Random-Forest-based Iterative Imputation-across Missing Completely at Random (MCAR), Missing at Random (MAR), and Missing Not at Random (MNAR) conditions. Three health datasets serve as experimental platforms: the Chronic Kidney Disease (CKD) dataset, the Heart Disease Dataset (HDD), and the Mice Protein Expression Dataset (MPED). Domain-knowledge-driven missingness patterns are constructed for each mechanism (Fig 2) at approximately 20-25% overall rates. Eight classifiers-KNN, Logistic Regression, SVC, Decision Tree, Random Forest, Gaussian Naïve Bayes, MLP, and a deep neural network-are trained on imputed data following GridSearchCV optimisation. Across all conditions, CHIT achieves near-perfect or perfect downstream classification, with SVC reaching up to 100% accuracy on the CKD dataset under MNAR, and 98.75% under MCAR and MAR. Competing methods degrade by up to 19.38 percentage points under MNAR, while CHIT's accuracy remains stable. Two structural properties account for this resilience: an iterative enrichment of the regression training set as records are completed, and a within-record prioritisation of the most data-scarce features for model-based filling. Taken together, these results provide mechanism-specific guidance for imputation selection in health informatics pipelines. Statistical significance of classifier accuracy differences between CHIT and MICE was assessed using McNemar's test (two-tailed, continuity-corrected chi-square) applied to the exact binary prediction vectors from each experiment [n_test = 80]. Ninety-five percent confidence intervals for all reported accuracy values were computed using the Wilson score method [z = 1.96]. The 42-configuration mean accuracy and standard deviation for CHIT are reported in Supplementary Table S1. Full McNemar results with confidence intervals are in Supplementary Table S2.
Related Concept Videos
Mechanistic Models: Compartment Models in Individual and Population Analysis
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Kaplan-Meier Approach
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Multiple Comparison Tests
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...