Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data01:16

Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data

Statistical inference techniques, paramount in hypothesis testing, differentiate into two broad categories: parametric and nonparametric statistics.
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance, comparing...
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving01:29

Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving

Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Weighted Mean00:57

Weighted Mean

While taking the arithmetic, geometric, or harmonic mean of a sample data set, equal importance is assigned to all the data points. However, all the values may not always be equally important in some data sets. An intrinsic bias might make it more important to give more weightage to specific values over others.
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
Introduction to Nonparametric Statistics01:28

Introduction to Nonparametric Statistics

Nonparametric statistics offer a powerful alternative to traditional parametric methods, useful when assumptions about the population distribution cannot be made. Unlike parametric tests, which require data to follow a specific distribution with well-defined parameters (such as the mean and standard deviation), nonparametric tests do not require such constraints. This makes them particularly valuable when dealing with small sample sizes, skewed data, or ordinal and categorical variables.
One of...
Residuals and Least-Squares Property01:11

Residuals and Least-Squares Property

The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Cluster Sampling Method01:20

Cluster Sampling Method

Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

NOTCH2NLC Intermediate-Length Repeat Expansions Are Associated with Parkinson Disease.

Annals of neurology·2020
Same author

AUTOPHAGY-RELATED14 and Its Associated Phosphatidylinositol 3-Kinase Complex Promote Autophagy in Arabidopsis.

The Plant cell·2020
Same author

Myocardial injury and risk factors for mortality in patients with COVID-19 pneumonia.

International journal of cardiology·2020
Same author

IDOL gene variant is associated with hyperlipidemia in Han population in Xinjiang, China.

Scientific reports·2020
Same author

LncRNA-5657 silencing alleviates sepsis-induced lung injury by suppressing the expression of spinster homology protein 2.

International immunopharmacology·2020
Same author

Hybrid Hydrogels for Synergistic Periodontal Antibacterial Treatment with Sustained Drug Release and NIR-Responsive Photothermal Effect.

International journal of nanomedicine·2020

Related Experiment Videos

A new dual-scale nearest neighbor statistical feature construction algorithm for imbalanced data oriented to Gaussian

Wei Wang1, Shuang Ouyang2, Fen Liu1

  • 1Business School, Guilin Tourism University, Guilin, 541006, China.

Scientific Reports
|June 13, 2026
PubMed
Summary

This study introduces a new feature construction algorithm, dynamic dual-scale nearest neighbor statistical ratio (NNDSR), to improve Gaussian Naive Bayes (GNB) performance on imbalanced datasets. NNDSR enhances minority class recognition and class separability, outperforming existing methods.

Keywords:
Dynamic dual-scale nearest neighbor statistical ratioFeature construction algorithmGaussian Naive Bayes classifierImbalanced datasets

Related Experiment Videos

Area of Science:

  • Machine Learning
  • Data Mining
  • Pattern Recognition

Background:

  • Gaussian Naive Bayes (GNB) classifier performance degrades on imbalanced datasets.
  • Sparse minority class features and severe class overlap contribute to GNB performance issues.
  • Existing methods like sampling and feature enhancement often mismatch GNB's core assumptions.

Purpose of the Study:

  • Propose a novel feature construction algorithm, dynamic dual-scale nearest neighbor statistical ratio (NNDSR), to address GNB performance degradation on imbalanced data.
  • Enhance the discriminability and Gaussian distribution adaptability of features.
  • Improve the recognition accuracy of minority classes and overall classification performance.

Main Methods:

  • Developed a dynamic dual-scale nearest neighbor mechanism to extract local aggregation and inter-class boundary information.
  • Generated new features using cross-class and dual-scale statistical ratio operations.
  • Validated the NNDSR algorithm on 22 UCI datasets with varying characteristics.

Main Results:

  • NNDSR significantly outperformed the original data and 16 mainstream algorithms across key metrics (AUC, G-mean, F-measure).
  • Demonstrated notable improvement in minority class recognition accuracy.
  • Confirmed efficiency and stability through scalability tests on large datasets.

Conclusions:

  • NNDSR is a robust feature construction algorithm for GNB to effectively handle imbalanced data.
  • The proposed method avoids information distortion issues of traditional sampling techniques.
  • NNDSR offers strong practical application value for imbalanced classification tasks.