Related Experiment Video
Updated: Nov 10, 2025

ExCYT: A Graphical User Interface for Streamlining Analysis of High-Dimensional Cytometry Data
Published on: January 16, 2019
Big data clustering techniques based on Spark: a literature review.
Mozamel M Saeed1, Zaher Al Aghbari2, Mohammed Alsharidah1
1Department of Computer Science, Prince Sattam Bin Abdul Aziz, Riyadh, Saudi Arabia.
This survey explores Apache Spark-based clustering methods for Big Data challenges. It introduces a new taxonomy and highlights future research directions in massive data clustering.
Area of Science:
- Data Mining
- Machine Learning
- Pattern Recognition
Background:
- Clustering is a key unsupervised learning technique challenged by massive data growth.
- Traditional methods struggle with Big Data, necessitating Big Data platform integration.
- Apache Spark offers fast, distributed processing for Big Data challenges.
Purpose of the Study:
- To systematically survey existing Apache Spark-based clustering methods.
- To evaluate these methods against Big Data characteristics.
- To propose a novel taxonomy for Spark-based clustering.
Main Methods:
- Systematic literature review of Spark-based clustering studies (2010-2020).
- Analysis of clustering methods concerning Big Data properties.
- Development of a new classification taxonomy.
Main Results:
- Identified and categorized existing Spark-based clustering approaches.
- Assessed the suitability of current methods for Big Data.
- Established a comprehensive overview of the field.
Conclusions:
- Spark-based clustering is an emerging research area with significant potential.
- A structured taxonomy is needed to understand and advance the field.
- Further research is required to address Big Data clustering challenges effectively.
Related Concept Videos
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Review and Preview
Scatter Plot
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Statistical Package for the Social Sciences (SPSS)
SPSS streamlines the process from data preparation to analysis and reporting. It is characterized by its user-friendly interface, which conceals...
Statistical Analysis: Overview
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...

