Related Experiment Video
Updated: Jul 2, 2026

Watershed Planning within a Quantitative Scenario Analysis Framework
Published on: July 24, 2016
A SMOTE PCA HDBSCAN approach for enhancing water quality classification in imbalanced datasets
Norashikin Nasaruddin1,2, Nurulkamal Masseran3, Wan Mohd Razi Idris4
1Department of Mathematical Sciences, Faculty of Science and Technology, Universiti Kebangsaan Malaysia, 43600, Bangi, Selangor, Malaysia. p119487@siswa.ukm.edu.my.
Abstract:
Class imbalance poses a significant challenge in water quality classification, often leading to biased predictions and diminished accuracy for minority classes. This study introduces SMOTE-PCA-HDBSCAN, a novel oversampling framework that integrates the Synthetic Minority Oversampling Technique (SMOTE) to generate synthetic samples, Principal Component Analysis (PCA) to enhance data separability, and Hierarchical Density-Based Spatial Clustering of Applications with Noise (HDBSCAN) to remove synthetic data noise. The cleaned synthetic data is then merged with the original dataset to form a balanced, noise-reduced training set. Comparative evaluations against SMOTE, SMOTE-DBSCAN, SMOTE-PCA-DBSCAN, SMOTE-ENN, and SMOTE-Tomek Links reveal that SMOTE-PCA-HDBSCAN consistently improves sensitivity for minority classes (Clean: 4.76% to 28.57%; Polluted: 38.09% to 61.90%) while maintaining high accuracy for the majority class. These results demonstrate the robustness of SMOTE-PCA-HDBSCAN in addressing class imbalance, offering a valuable tool for enhancing predictive models in environmental monitoring and other domains with imbalanced datasets.
Related Concept Videos
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Sampling Plans
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...

