Related Experiment Video
Updated: Sep 3, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Addressing Binary Classification over Class Imbalanced Clinical Datasets Using Computationally Intelligent Techniques
Vinod Kumar1, Gotam Singh Lalotra2, Ponnusamy Sasikala3
1Computer Science and Engineering, Koneru Lakshmaiah Education Foundation, Vaddeswaram 522302, India.
This study evaluated six classifiers and seven balancing techniques on imbalanced clinical data. SMOTEEN demonstrated superior performance for improving machine learning model accuracy in healthcare applications.
Area of Science:
- Machine Learning in Healthcare
- Data Science
- Clinical Informatics
Background:
- Clinical datasets are crucial for intelligent healthcare systems.
- Real-world clinical data often exhibits class imbalance, leading to poor classifier performance.
- This imbalance negatively impacts accuracy, precision, and recall, necessitating effective data balancing techniques.
Purpose of the Study:
- To empirically evaluate the performance of six classifiers on five imbalanced clinical datasets.
- To compare seven distinct class balancing techniques for addressing data imbalance in clinical machine learning.
- To identify optimal strategies for handling imbalanced clinical datasets in supervised learning.
Main Methods:
- Literature review on class imbalanced learning.
- Performance evaluation of Decision Tree, k-Nearest Neighbor, Logistic Regression, Artificial Neural Network, Support Vector Machine, and Gaussian Naïve Bayes classifiers.
- Application and comparison of seven balancing techniques: Undersampling, Random Oversampling, SMOTE, ADASYN, SVM-SMOTE, SMOTEEN, and SMOTETOMEK.
Main Results:
- SMOTEEN consistently outperformed the other six data-balancing techniques across all tested classifiers and datasets.
- The remaining six balancing techniques showed comparable, yet moderately lesser, performance compared to SMOTEEN.
- Analysis explored the reasons behind the effectiveness of specific classifiers and balancing methods.
Conclusions:
- SMOTEEN is a highly effective technique for addressing class imbalance in clinical datasets.
- The findings provide practical recommendations for improving supervised machine learning model performance in healthcare.
- Careful selection of data-balancing techniques is essential for robust clinical predictive models.
More Related Videos
12:18A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
Related Concept Videos
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Classification of Systems-I
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Classification of Systems-II
Aggregates Classification
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Classification of Leukocytes
Neutrophils are the most abundant type of granular leukocytes, comprising 50-70% of all leukocytes. They feature small, evenly distributed granules and a...
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...