Related Experiment Video
Updated: Oct 25, 2025

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Measuring the Effectiveness of Adaptive Random Forest for Handling Concept Drift in Big Data Streams
Abdulaziz O AlQabbany1,2, Aqil M Azmi1
1Department of Computer Science, College of Computer & Information Sciences, King Saud University, Riyadh 11543, Saudi Arabia.
This study enhances Adaptive Random Forest (ARF) for big data stream processing by optimizing resampling effectiveness (ρ). Tuning the Poisson distribution parameter (λ) improves accuracy and execution time, crucial for real-time big data analytics.
Area of Science:
- Machine Learning
- Data Science
- Stream Data Processing
Background:
- Big data and stream data processing present real-time challenges.
- Concept drift significantly impacts machine learning models on data streams.
- Adaptive Random Forest (ARF) is a promising stream learning algorithm for handling data distribution changes.
Purpose of the Study:
- To propose and evaluate a resampling mechanism for enhancing the efficiency of streaming algorithms like ARF.
- To introduce a measure of resampling effectiveness (ρ) that balances accuracy and execution time.
- To optimize ARF performance by tuning the Poisson distribution parameter (λ) for resampling.
Main Methods:
- Empirical selection of the Poisson distribution parameter (λ) using six synthetic datasets with varying drift types.
- Comparison of standard ARF with tuned variations based on resampling effectiveness (ρ).
- Validation of the proposed enhancement method on three real-world case studies: Amazon reviews, Arabic hotel reviews, and COVID-19 tweet sentiment analysis.
Main Results:
- The proposed resampling enhancement method demonstrated considerable improvement in ARF performance across various scenarios.
- Tuning the Poisson distribution parameter (λ) effectively balances accuracy and execution time for online learning.
- The method proved effective in processing large datasets and handling different types of concept drift.
Conclusions:
- Optimizing resampling strategies is key to improving the efficiency of stream learning algorithms.
- The proposed method offers a practical enhancement for ARF, leading to better performance in big data applications.
- This approach is effective for real-time processing of diverse data streams, including sentiment analysis.
Related Concept Videos
Survival Tree
Building a Survival Tree
Constructing a...
Genetic Drift
Randomized Experiments
Simple randomization
Simple...
Instinctive Drift
Regression Toward the Mean
Random and Systematic Errors

