Related Experiment Video
Updated: Aug 12, 2025

Author Spotlight: AI-Driven Trypanosome Species Detection from Microscopic Images
Published on: October 27, 2023
Analyzing the effect of data preprocessing techniques using machine learning algorithms on the diagnosis of COVID-19
Gizemnur Erol1, Betül Uzbaş2, Cüneyt Yücelbaş3
1Konya Technical University Software Engineering Department Konya Turkey.
Machine learning models for COVID-19 detection using blood data show improved accuracy with specific data preprocessing techniques. Optimizing these methods is key for reliable diagnostic tools.
Area of Science:
- Medical diagnostics
- Machine learning in healthcare
- Data science
Background:
- Real-time polymerase chain reaction (RT-PCR) tests are insufficient for rapid COVID-19 diagnosis due to global spread.
- Medical data, including COVID-19 blood parameters, presents challenges like inconsistency, incompleteness, and large scale.
- Limited and potentially biased medical datasets can adversely affect machine learning study accuracy.
Purpose of the Study:
- To investigate the impact of various data preprocessing techniques on the classification accuracy of COVID-19 using blood parameters.
- To improve existing, limited datasets of COVID-19 blood data for more consistent machine learning results.
- To compare the effectiveness of different imputation and balancing methods in enhancing COVID-19 classification models.
Main Methods:
- Applied feature encoding and scaling to a dataset of 279 patients' blood data.
- Utilized K-nearest neighbor (KNN) imputation and Multiple Imputation by Chained Equations (MICE) to handle missing data.
- Employed Synthetic Minority Over-sampling Technique (SMOTE) for data balancing.
- Evaluated the performance of ensemble (bagging, AdaBoost, random forest) and popular classifier algorithms (KNN, SVM, logistic regression, ANN, decision tree).
Main Results:
- The highest accuracy achieved with the bagging classifier was 83.91% using KNN imputation without SMOTE.
- With SMOTE, the bagging classifier achieved accuracies of 83.42% (KNN imputation) and 83.74% (MICE imputation).
- Comparative analysis demonstrated that specific data preprocessing combinations significantly influence classification success.
Conclusions:
- Data preprocessing techniques critically impact the accuracy of machine learning models for COVID-19 classification from blood data.
- The choice of imputation (KNN vs. MICE) and balancing (SMOTE) methods affects model performance.
- Experimental results highlight the importance of selecting the right combination of preprocessing steps for successful COVID-19 diagnostic models.
More Related Videos
07:15Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
Related Concept Videos
Steps in Outbreak Investigation
Issues And Trends In Healthcare Delivery System
Cost Containment
Payment for healthcare services has historically promoted adoption of costly and often unnecessary or inefficient...