Related Experiment Video
Updated: May 7, 2026

SIVQ-LCM Protocol for the ArcturusXT Instrument
Published on: July 23, 2014
LVQ-SMOTE - Learning Vector Quantization based Synthetic Minority Over-sampling Technique for biomedical data
Munehiro Nakamura1, Yusuke Kajiwara, Atsushi Otsuka
1Department of Natural Science and Engineering, Kanazawa University, Ishikawa 9200941, Japan. m-nakamura@blitz.ec.t.kanazawa-u.ac.jp.
This study introduces a novel over-sampling method using learning vector quantization codebooks to improve classification of imbalanced biomedical data. The new Synthetic Minority Over-sampling Technique (SMOTE) generates more effective synthetic samples, enhancing model performance.
Area of Science:
- Biomedical Data Science
- Machine Learning
- Bioinformatics
Background:
- Imbalanced biomedical datasets pose challenges for traditional classification algorithms.
- Existing Synthetic Minority Over-sampling Technique (SMOTE) methods offer limited improvements over basic SMOTE.
- Vast empty feature spaces hinder accurate borderline estimation between classes.
Purpose of the Study:
- To develop a novel over-sampling method to enhance the classification of imbalanced biomedical data.
- To generate synthetic samples that occupy more feature space, improving upon existing SMOTE algorithms.
- To leverage learning vector quantization codebooks for more effective synthetic data generation.
Main Methods:
- A new over-sampling method was developed utilizing codebooks from learning vector quantization.
- The proposed method generates synthetic samples by referencing actual data points.
- Integration with existing SMOTE variants, such as MWMOTE, was explored.
Main Results:
- The novel over-sampling method demonstrated superior performance compared to basic SMOTE on four out of five standard classification algorithms across eight real-world imbalanced datasets.
- Performance gains were observed when the proposed method was combined with MWMOTE (Minority-over-sampling Weighted Majority Technique).
- Analysis of β-turn types prediction datasets revealed novel patterns previously unobserved.
Conclusions:
- The proposed over-sampling method effectively generates useful synthetic samples for imbalanced biomedical data classification.
- The method is compatible with standard classification algorithms and existing over-sampling techniques.
- This approach offers a promising solution for improving machine learning model accuracy in biomedical applications.
Related Concept Videos
Sampling Methods: Overview
In analytical chemistry, the choice of sampling...
Upsampling
Sampling Methods: Sample Types
Solid samples include a variety of substances, such as sediments from water bodies, soil, metals, and biological tissues. Two standard methods for extracting sediments from water bodies are grab sampling and piston coring. Grab sampling involves using a device to collect a discrete sediment sample from the bottom of a water body with minimal disturbance. Grab samples do not always represent the entire area due to...
Improving Translational Accuracy
Improving Translational Accuracy
Sampling Continuous Time Signal
In the...
