Related Experiment Video
Updated: Jan 21, 2026

Patient-specific Modeling of the Heart: Estimation of Ventricular Fiber Orientations
Published on: January 8, 2013
Estimating the success of re-identifications in incomplete datasets using generative models
Luc Rocher1,2,3, Julien M Hendrickx1, Yves-Alexandre de Montjoye4,5
1Information and Communication Technologies, Electronics and Applied Mathematics (ICTEAM), Université catholique de Louvain, B-1348, Louvain-la-Neuve, Belgium.
Modern de-identification methods fail to protect privacy. Even with incomplete data, our generative copula model accurately predicts re-identification risk, showing most anonymized datasets are not GDPR-compliant.
Area of Science:
- Data privacy
- Statistical modeling
- Computational social science
Background:
- Rich datasets are crucial for data-driven research but pose privacy risks.
- De-identification and sampling are standard methods to anonymize data.
- Existing anonymization techniques may not meet current privacy standards.
Purpose of the Study:
- To develop a generative copula-based method for estimating individual re-identification likelihood.
- To assess the effectiveness of de-identification in protecting privacy in large datasets.
- To evaluate whether current anonymization practices align with GDPR standards.
Main Methods:
- Utilized a generative copula-based approach to model data distributions.
- Estimated the probability of correct re-identification for individuals within datasets.
- Validated the method on 210 diverse populations, assessing predictive accuracy using AUC scores.
Main Results:
- The proposed method achieved high AUC scores (0.84-0.97) for predicting individual uniqueness.
- A low false-discovery rate was observed, indicating reliable re-identification risk estimation.
- Analysis revealed that 99.98% of Americans could be re-identified using 15 demographic attributes.
Conclusions:
- Current de-identification practices, even with heavy sampling, are insufficient for modern privacy standards like GDPR.
- The 'release-and-forget' model of data anonymization is technically and legally inadequate.
- There is a critical need for more robust privacy-preserving techniques in data sharing.
Related Concept Videos
Ecological Succession
Incomplete Dominance
What are Estimates?
The estimate for the mean of a sample is denoted by ͞x, whereas the mean of the population is designated as μ. Further, parameters such...
One-Compartment Open Model for IV Bolus Administration: Estimation of Clearance
In the one-compartment open model for intravenous (IV) bolus administration, clearance is estimated by dividing the elimination rate by the plasma drug concentration. This equation leverages the elimination rate constant and the apparent...
Estimation of k and VD of Aminoglycosides
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
On...

