Challenges in Preprocessing Routine Laboratory Data for Machine Learning
Katharina Wendt1, Michael Marschollek1, Thomas Illig2
1Peter L. Reichertz Institute for Medical Informatics of TU Braunschweig and Hannover Medical School, Hannover Medical School, Hannover, Germany.
None:
Post-Covid syndrome remains a major clinical challenge due to the lack of specific biomarkers and reliance on exclusion criteria. Routine laboratory data offers potential for identifying biological signatures, but their use in machine learning is hampered by data quality issues. This study investigates the impact of preprocessing on classification performance using 52 laboratory parameters from the German NAPKON cohort (n = 1,292: 1,130 Covid recovered, 162 post-Covid patients), measured at four time points. Data preprocessing included unit harmonization, handling of missing values and statistical assessment of inter-laboratory variability. Non-parametric tests (Kruskal-Wallis, Wilcoxon) revealed significant differences in laboratory values across units (p < 0.05), even after harmonization. The best classification performance was achieved at 70% allowed missingness per sample and feature. Our findings underscore that handling missing values and data harmonization are particularly crucial and have a major impact on model performance. But even after preprocessing, residual variability persists due to biological and technical factors.
