Related Experiment Video
Updated: Aug 1, 2025

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
A self-inspected adaptive SMOTE algorithm (SASMOTE) for highly imbalanced data classification in healthcare.
Tanapol Kosolwattana1, Chenang Liu2, Renjie Hu3
1Department of Industrial Engineering, University of Houston, Houston, USA.
A new Self-Inspected Adaptive SMOTE (SASMOTE) model improves imbalanced healthcare data classification by generating higher-quality synthetic samples. This approach enhances machine learning model usability for rare disease prediction and risk gene discovery.
Area of Science:
- Machine Learning
- Bioinformatics
- Medical Informatics
Background:
- Healthcare datasets often exhibit class imbalance due to rare events like disease onset.
- The Synthetic Minority Over-sampling Technique (SMOTE) is used for imbalanced data but can generate low-quality, ambiguous samples.
- Existing SMOTE methods struggle with sample quality, impacting classification performance in critical healthcare applications.
Purpose of the Study:
- To introduce a novel Self-Inspected Adaptive SMOTE (SASMOTE) model for enhancing synthetic sample quality in imbalanced datasets.
- To improve the accuracy and reliability of machine learning models in healthcare by addressing data imbalance.
- To develop a method that generates higher-quality, separable synthetic minority class samples.
Main Methods:
- Proposed SASMOTE model utilizing an adaptive nearest neighborhood selection algorithm.
- Implemented an uncertainty elimination via self-inspection approach to filter ambiguous generated samples.
- Compared SASMOTE with existing SMOTE-based algorithms on real-world healthcare datasets.
Main Results:
- SASMOTE generates higher-quality synthetic samples compared to traditional SMOTE algorithms.
- Demonstrated improved prediction performance, particularly in F1 score, across two healthcare case studies.
- Effectively addressed challenges of ambiguous and non-separable samples in imbalanced data.
Conclusions:
- The proposed SASMOTE model significantly enhances the quality of synthetic samples for imbalanced healthcare data.
- SASMOTE offers a promising solution for improving machine learning model performance in critical applications like risk gene discovery and disease prediction.
- The self-inspection and adaptive neighborhood selection mechanisms contribute to better classification outcomes on highly imbalanced datasets.
More Related Videos
06:55Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
07:35Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Related Concept Videos
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Documentation in Long-Term and Home Healthcare Setting
Long-Term Care Facilities
Aggregates Classification
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Integrated Healthcare System