Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Strategies for Assessing and Addressing Confounding01:25

Strategies for Assessing and Addressing Confounding

189
Confounding is a critical issue in epidemiological studies, often leading to misleading conclusions about associations between exposures and outcomes. It occurs when the relationship between the exposure and the outcome is mixed with the effects of other factors that influence the outcome. Given that, addressing confounding is of high importance for drawing accurate inferences in research.
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
189
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches01:23

Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches

216
Biopharmaceutical studies constitute a vital field aiming to enhance drug delivery methods and refine therapeutic approaches, drawing upon diverse interdisciplinary knowledge. In research methodologies, the choice between controlled and non-controlled studies significantly influences the study's reliability and accuracy.
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
216
Regression Toward the Mean01:52

Regression Toward the Mean

6.6K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.6K
Kaplan-Meier Approach01:24

Kaplan-Meier Approach

340
The Kaplan-Meier estimator is a non-parametric method used to estimate the survival function from time-to-event data. In medical research, it is frequently employed to measure the proportion of patients surviving for a certain period after treatment. This estimator is fundamental in analyzing time-to-event data, making it indispensable in clinical trials, epidemiological studies, and reliability engineering. By estimating survival probabilities, researchers can evaluate treatment effectiveness,...
340
Statistical Methods for Analyzing Epidemiological Data01:25

Statistical Methods for Analyzing Epidemiological Data

650
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
650
Comparing the Survival Analysis of Two or More Groups01:20

Comparing the Survival Analysis of Two or More Groups

369
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
369

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

An intelligent autoregressive-distributed lag model: A climate-driven approach for predicting dengue fever incidence in Taiwan cities.

Acta tropica·2025
Same author

A streamlined U-Net convolution network for medical image processing.

Quantitative imaging in medicine and surgery·2025
Same author

The relationship between hotel star rating and website information quality based on visual presentation.

PloS one·2023
Same author

Synthesis of Quercetin-Acid Esters and Its Reduction of H<sub>2</sub> O<sub>2</sub> -Triggered PC12 Cells Damage by Down-Regulating ROS.

Chemistry & biodiversity·2023
Same author

Rule-based classifier based on accident frequency and three-stage dimensionality reduction for exploring the factors of road accident injuries.

PloS one·2022
Same author

Identification of the Framingham Risk Score by an Entropy-Based Rule Model for Cardiovascular Disease.

Entropy (Basel, Switzerland)·2020

Related Experiment Video

Updated: Nov 3, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
06:55

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index

Published on: January 8, 2020

14.8K

A multiple combined method for rebalancing medical data with class imbalances.

Yun-Chun Wang1, Ching-Hsue Cheng1

  • 1Department of Information Management, National Yunlin University of Science & Technology, Touliou, Yunlin, 640, Taiwan.

Computers in Biology and Medicine
|June 6, 2021
PubMed
Summary

This study introduces a combined method to rebalance imbalanced medical data, improving classification accuracy for minority classes. The approach enhances model performance using resampling, optimization, and cost-sensitive learning for critical medical diagnoses.

Keywords:
Class imbalanceMetaCostParticle swarm optimizationSynthetic minority oversampling technique

More Related Videos

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.7K
A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
12:18

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment

Published on: January 11, 2020

7.7K

Related Experiment Videos

Last Updated: Nov 3, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
06:55

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index

Published on: January 8, 2020

14.8K
Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.7K
A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
12:18

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment

Published on: January 11, 2020

7.7K

Area of Science:

  • Medical Informatics
  • Machine Learning
  • Data Science

Background:

  • Medical datasets frequently exhibit class imbalance, negatively impacting classification model performance and minority class accuracy.
  • Accurate identification is critical in medicine due to the high cost of misclassification and potential patient harm.

Purpose of the Study:

  • To propose and validate a novel, multiple combined method for rebalancing imbalanced medical data.
  • To enhance the accuracy and reliability of classification models in critical medical applications.

Main Methods:

  • A hybrid approach combining resampling techniques (Synthetic Minority Oversampling Technique [SMOTE], Undersampling [US]), Particle Swarm Optimization (PSO), and MetaCost.
  • Experimental validation using nine medical datasets and decision tree analysis for rule generation.
  • Comparison against existing methods to evaluate performance improvements.

Main Results:

  • The proposed ensemble learning method significantly improved Area Under the ROC Curve (AUC), recall, precision, and F1-score.
  • MetaCost enhanced sensitivity, SMOTE boosted AUC, and US improved sensitivity, F1-score, and reduced misclassification costs in highly imbalanced data.
  • PSO-based attribute selection increased sensitivity and reduced data dimensionality.

Conclusions:

  • The combined method effectively addresses class imbalance in medical data, leading to more reliable diagnostic models.
  • Specific strategies (US for imbalance ratio >9, combined SMOTE/US for <9) are recommended based on imbalance levels.
  • The findings support the use of advanced rebalancing techniques for improving patient outcomes in data-driven healthcare.