Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Sampling Methods: Overview01:06

Sampling Methods: Overview

269
A sample refers to a smaller subset representative of a larger population. In analytical chemistry, studying or analyzing an entire population is often impractical or impossible. Therefore, samples are used to draw inferences and generalize the whole population. The sampling method selects individuals or items from a population to create a sample. Standard sampling methods include random, judgemental, systematic, stratified, and cluster sampling. 
In analytical chemistry, the choice of...
269
Improving Translational Accuracy02:07

Improving Translational Accuracy

8.5K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
8.5K
Downsampling01:20

Downsampling

126
When considering a sampled sequence with zero values between sampling instants, one can replace it by taking every N-th value of the sequence. At these integer multiples of N, the original and sampled sequences coincide. This process, known as decimation, involves extracting every N-th sample from a sequence, thereby creating a more efficient sequence.
The Fourier transform of the decimated sequence reveals a combination of scaled and shifted versions of the original spectrum. This...
126
Systematic Sampling Method01:17

Systematic Sampling Method

9.9K
Sampling is a technique to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. Data are the result of sampling from a population. The sampling method ensures that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
Systematic sampling is one of the simplest methods...
9.9K
Bias in Epidemiological Studies01:29

Bias in Epidemiological Studies

131
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:  
131
Random Sampling Method01:09

Random Sampling Method

10.9K
Sampling is a technique to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. Data are the result of sampling from a population. The sampling method ensures that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest. Among the various sampling methods used by...
10.9K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Disparities in the Uptake of COVID-19 Vaccination Between Māori and Non-Māori in Aotearoa New Zealand.

Journal of the Royal Society of New Zealand·2026
Same author

MEHC-Curation: A Python Framework for High-Quality Molecular Data Set Curation.

Journal of chemical information and modeling·2026
Same author

Edge-updating graph neural networks for modeling feature interactions in tabular data.

Neural networks : the official journal of the International Neural Network Society·2026
Same author

Emerging Artificial Intelligence Methodologies in Computational Biology.

Journal of molecular biology·2025
Same author

Masked graph transformer for blood-brain barrier permeability prediction.

Journal of molecular biology·2025
Same author

Deterministic Autoencoder using Wasserstein loss for tabular data generation.

Neural networks : the official journal of the International Neural Network Society·2025

Related Experiment Video

Updated: May 26, 2025

Author Spotlight: Advancing Alzheimer's Research – Exploring Early Detection and Multi-Omics Approaches
09:47

Author Spotlight: Advancing Alzheimer's Research – Exploring Early Detection and Multi-Omics Approaches

Published on: December 15, 2023

940

Addressing imbalance in health data: Synthetic minority oversampling using deep learning.

Alex X Wang1, Viet-Tuan Le2, Hau Nguyen Trung2

  • 1School of Mathematics and Statistics, Victoria University of Wellington, Kelburn Parade, Wellington 6012, New Zealand.

Computers in Biology and Medicine
|February 21, 2025
PubMed
Summary

This study introduces an advanced deep learning method to tackle class imbalance in healthcare data, improving machine learning model fairness and patient safety by generating synthetic positive samples and refining majority class data.

Keywords:
Contrastive learningDeep learningImbalance dataSynthetic data

More Related Videos

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
06:55

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index

Published on: January 8, 2020

14.3K
A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
12:18

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment

Published on: January 11, 2020

7.4K

Related Experiment Videos

Last Updated: May 26, 2025

Author Spotlight: Advancing Alzheimer's Research – Exploring Early Detection and Multi-Omics Approaches
09:47

Author Spotlight: Advancing Alzheimer's Research – Exploring Early Detection and Multi-Omics Approaches

Published on: December 15, 2023

940
Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
06:55

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index

Published on: January 8, 2020

14.3K
A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
12:18

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment

Published on: January 11, 2020

7.4K

Area of Science:

  • Machine Learning
  • Artificial Intelligence
  • Healthcare Informatics

Background:

  • Class imbalance in healthcare data leads to biased machine learning models, compromising patient safety and healthcare delivery.
  • Traditional oversampling methods like SMOTE have limitations in handling complex data, heterogeneous types, and multi-class scenarios.

Purpose of the Study:

  • To propose a novel deep learning approach for addressing class imbalance in healthcare datasets.
  • To enhance the performance and fairness of machine learning models in clinical applications.

Main Methods:

  • An Auxiliary-guided Conditional Variational Autoencoder (ACVAE) with contrastive learning was developed for synthetic data generation.
  • An ensemble technique combining ACVAE for oversampling positive cases and Edited Centroid-Displacement Nearest Neighbor (ECDNN) for majority class undersampling was employed.

Main Results:

  • Experiments on 12 health datasets demonstrated the effectiveness of the proposed ACVAE-ECDNN ensemble method.
  • The approach showed significant improvements in model performance across various metrics compared to traditional oversampling techniques.

Conclusions:

  • Deep learning-based synthetic oversampling offers a powerful solution for class imbalance in healthcare data.
  • The proposed method enhances dataset balance and informativeness, leading to more reliable machine learning models for healthcare.