Utility-driven assessment of anonymized data via clustering
Maria Eugénia Ferrão1, Paula Prata2, Paulo Fazendeiro3
1Universidade da Beira Interior, Covilha, Portugal and CEMAPRE, Lisboa, Portugal.
Scientific Data
|July 30, 2022
Summary
Data anonymization techniques, like k-anonymity and differential privacy, can compromise clustering analysis. This study found that anonymizing low-dimensionality datasets biases field-of-study estimates for law students.
Area of Science:
- Data Science
- Social Sciences
Background:
- Clustering is valuable for identifying groups of interest in datasets.
- Data anonymization is crucial for protecting privacy but can impact data utility.
- Assessing the utility of anonymized data for analytical tasks like clustering is essential.
Purpose of the Study:
- To explore clustering techniques as data utility models within data anonymization frameworks.
- To evaluate the impact of anonymization (k-anonymity, differential privacy) on clustering results.
- To assess the utility of anonymized data using metrics, group characteristics, and relative risk.
Main Methods:
- Applied partitional clustering to a dataset of Portuguese higher education Law students.
- Compared anonymized clustering scenarios against the original data.
- Utilized clustering validity indices and standard metrics to evaluate data utility and structure preservation.
- Examined relative risk as a metric for social sciences research.
Main Results:
- For low dimensionality/cardinality datasets, anonymization procedures significantly jeopardize clustering.
- Evidence suggests that anonymized data leads to biased estimates of field-of-study.
- The effectiveness of clustering as a data utility model in anonymization is challenged by data characteristics.
Conclusions:
- Data anonymization can severely impact the integrity of clustering analyses, particularly for smaller datasets.
- The utility of anonymized data for inferring group characteristics, such as field of study, is questionable.
- Careful consideration of anonymization methods and dataset properties is necessary to maintain analytical validity.
Related Concept Videos
Cluster Sampling Method
12.5K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
12.5K
Sampling Plans
253
Sampling is a crucial step in analytical chemistry, allowing researchers to collect representative data from a large population. Common sampling methods include random, judgmental, systematic, stratified, and cluster sampling.
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
253
Model-Independent Approaches for Pharmacokinetic Data: Noncompartmental Analysis
115
Noncompartmental analyses offer an alternative method for describing drug pharmacokinetics without relying on a specific compartmental model. In this approach, the drug's pharmacokinetics are assumed to be linear, with the terminal phase log-linear. This assumption allows for simplified analysis and interpretation of the drug's behavior in the body.
One important characteristic of noncompartmental analyses is that drug exposure increases proportionally with increasing doses. This...
One important characteristic of noncompartmental analyses is that drug exposure increases proportionally with increasing doses. This...
115
Statistical Methods to Analyze Parametric Data: ANOVA
640
Analysis of Variance, or ANOVA, is a powerful statistical technique used to analyze parametric data, primarily in research and experimental studies. It's designed to compare the means of two or more groups, assisting researchers in identifying any significant differences between these group means. There are two main types of ANOVA based on the complexity of the analysis: one-way and two-way.
One-way ANOVA is applied when a single independent variable or factor is scrutinized. It compares...
One-way ANOVA is applied when a single independent variable or factor is scrutinized. It compares...
640


