Related Experiment Video
Updated: Aug 10, 2026

07:35
Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Classification of Imbalanced Data by Oversampling in Kernel Space of Support Vector Machines
IEEE Transactions on Neural Networks and Learning Systems
|October 14, 2017
Summary
This study introduces weighted kernel-based SMOTE (WK-SMOTE) to improve fault diagnosis in imbalanced industrial machine data. The novel method enhances classifier performance on complex, imbalanced datasets, aiding in early equipment failure detection.
Area of Science:
- Machine Learning
- Industrial Engineering
- Data Science
Background:
- Industrial machine fault diagnosis data is often imbalanced, posing challenges for traditional classifiers.
- Imbalanced datasets lead to biased models, hindering accurate fault stage diagnosis.
- Existing methods like synthetic minority oversampling technique (SMOTE) have limitations with nonlinear problems.
Purpose of the Study:
- To develop an improved oversampling technique for imbalanced datasets in industrial fault diagnosis.
- To address the limitations of traditional SMOTE in nonlinear feature spaces.
- To create a hierarchical framework for multiclass imbalanced problems with progressive class orders.
Main Methods:
- Proposing weighted kernel-based SMOTE (WK-SMOTE) for oversampling in the feature space of Support Vector Machine (SVM) classifiers.
- Integrating WK-SMOTE with a cost-sensitive SVM formulation.
- Developing a hierarchical framework for multiclass imbalanced problems with ordered classes.
Main Results:
- WK-SMOTE combined with cost-sensitive SVM significantly improves performance on benchmark imbalanced datasets.
- The proposed methods outperform baseline approaches in fault diagnosis tasks.
- Validation on a real-world industrial problem demonstrated effectiveness in detecting high-voltage equipment insulation deterioration.
Conclusions:
- WK-SMOTE offers a robust solution for imbalanced data challenges in industrial fault diagnosis.
- The hierarchical framework effectively handles multiclass imbalanced problems with progressive class structures.
- The developed techniques enhance the reliability and accuracy of industrial equipment monitoring systems.
Related Concept Videos
Cluster Sampling Method
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
One-Way ANOVA: Unequal Sample Sizes
One-way ANOVA can be performed on three or more samples of unequal sizes. However, calculations get complicated when sample sizes are not always the same. So, while performing ANOVA with unequal samples size, the following equation is used:
Quantifying and Rejecting Outliers: The Grubbs Test
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This number is...
Sampling Methods: Overview
A sample refers to a smaller subset representative of a larger population. In analytical chemistry, studying or analyzing an entire population is often impractical or impossible. Therefore, samples are used to draw inferences and generalize the whole population. The sampling method selects individuals or items from a population to create a sample. Standard sampling methods include random, judgemental, systematic, stratified, and cluster sampling.
In analytical chemistry, the choice of sampling...
In analytical chemistry, the choice of sampling...
Classification of Signals
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Aggregates Classification
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
