Related Experiment Video
Updated: Jan 11, 2026

03:37
Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
Published on: March 1, 2024
1.2K
Hybrid Synthetic Minority Over-sampling Technique (HSMOTE) and Ensemble Deep Dynamic Classifier Model (EDDCM) for big
Priyadharsini M1, Bhawana Tyagi2, Naga Priyadarsini R1
1School of Computer Science and Engineering, Vellore Institute of Technology, Vellore, Tamilnadu, India.
Scientific Reports
|November 11, 2025
Summary
This study introduces a hybrid framework to improve Big Data Classification (BDC) by addressing class imbalance and high dimensionality. The novel approach enhances machine learning model performance for critical applications.
Area of Science:
- Computer Science
- Machine Learning
- Data Science
Background:
- Big Data Classification (BDC) is crucial in healthcare, e-commerce, and banking.
- Conventional machine learning models struggle with high dimensionality and class imbalance in BDC.
- Existing methods often fail to effectively handle imbalanced datasets and select relevant features.
Purpose of the Study:
- To propose a hybrid framework enhancing Big Data Classification (BDC) effectiveness.
- To address challenges of high dimensionality and class imbalance in BDC.
- To improve the reliability and accuracy of classification models in complex datasets.
Main Methods:
- Introduced Hybrid Synthetic Minority Over-sampling Technique (HSMOTE) for class imbalance handling.
- Developed Optimization Ensemble Feature Selection Model (OEFSM) using meta-heuristic algorithms (FWDFA, AEHO, FWGWO) for robust feature selection.
- Proposed Ensemble Deep Dynamic Classifier Model (EDDCM) integrating DWCNN, DWBi-LSTM, and WAE with a dynamic ensemble strategy.
Main Results:
- The proposed framework demonstrated improved classification results across various datasets.
- Significant performance enhancements were observed particularly under conditions of class imbalance and high dimensionality.
- The integration of HSMOTE, OEFSM, and EDDCM effectively improved precision, recall, F-measure, and accuracy.
Conclusions:
- The hybrid framework offers a robust solution for Big Data Classification challenges.
- The combination of advanced sampling, feature selection, and deep learning ensemble methods enhances model performance.
- This approach provides a reliable method for improving classification accuracy in imbalanced and high-dimensional data.
Related Concept Videos
Aggregates Classification
960
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
960
Classification of Systems-I
543
Linearity is a system property characterized by a direct input-output relationship, combining homogeneity and additivity.
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
543
Classification of Systems-II
449
Continuous-time systems have continuous input and output signals, with time measured continuously. These systems are generally defined by differential or algebraic equations. For instance, in an RC circuit, the relationship between input and output voltage is expressed through a differential equation derived from Ohm's law and the capacitor relation,
449
Cluster Sampling Method
13.9K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
13.9K
Classification of Signals
1.3K
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
1.3K

