Related Experiment Video
Updated: Aug 1, 2025

Large Scale Energy Efficient Sensor Network Routing Using a Quantum Processor Unit
Published on: September 8, 2023
A distributed computing model for big data anonymization in the networks.
Farough Ashkouti1, Keyhan Khamforoosh2
1Department of Computer Engineering, Mahabad Branch, Islamic Azad University, Mahabad, Iran.
This study introduces a novel Apache Spark-based model for anonymizing big data, ensuring privacy during data publishing. It efficiently preserves λ-diversity using in-memory computations and RDD programming for scalable, robust data analysis.
Area of Science:
- Computer Science
- Data Science
- Information Security
Background:
- Big data's rapid growth presents significant challenges for IT infrastructure and computing capacity.
- Publishing big data for analysis risks individual privacy, necessitating robust anonymization techniques.
- Apache Spark offers a scalable, in-memory computing framework ideal for large-scale data processing.
Purpose of the Study:
- To propose an efficient parallel computing model for privacy-preserving big data anonymization.
- To leverage Apache Spark's resilient distributed dataset (RDD) programming for enhanced data anonymization.
- To address runtime, scalability, and performance issues in large-scale data anonymization.
Main Methods:
- Developed a three-phase in-memory computation model for big data anonymization using Apache Spark.
- Implemented partition-based data clustering algorithms to support the λ-diversity privacy model.
- Utilized RDD transformations and actions, incorporating City block and Pearson distance functions.
Main Results:
- Achieved efficient parallel implementation of a novel big data anonymization computing model.
- Demonstrated the model's effectiveness in preserving the λ-diversity privacy model.
- Provided a comprehensive guideline for applying Apache Spark in privacy-preserving big data research.
Conclusions:
- Apache Spark is a suitable framework for implementing scalable and performant privacy-preserving big data anonymization.
- The proposed three-phase in-memory model effectively handles the complexities of large-scale data anonymization.
- The Spark-based implementation offers a valuable tool for researchers in the field of data privacy.
More Related Videos
08:53Integrating Computerized Linguistic and Social Network Analyses to Capture Addiction Recovery Capital in an Online Community
Published on: May 31, 2019
09:47Author Spotlight: Advancing Alzheimer's Research – Exploring Early Detection and Multi-Omics Approaches
Published on: December 15, 2023
Related Concept Videos
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Model-Independent Approaches for Pharmacokinetic Data: Noncompartmental Analysis
One important characteristic of noncompartmental analyses is that drug exposure increases proportionally with increasing doses. This...
Analysis of Population Pharmacokinetic Data
Deindividuation
Analysis Methods of Pharmacokinetic Data: Model and Model-Independent Approaches
The model approach uses mathematical models to describe changes in drug concentration over time. Pharmacokinetic models help characterize drug behavior in patients, predict drug concentration in the body fluids, calculate optimum dosage regimens, and evaluate the risk of toxicity. However, ensuring that the model fits the experimental data accurately...
Censoring Survival Data