Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Videos

Effects of data anonymization by cell suppression on descriptive statistics and predictive modeling performance.

L Ohno-Machado1, S A Vinterbo, S Dreiseitl

  • 1Decision Systems Group, Brigham & Women's Hospital, Harvard Medical School, Boston, MA, USA.

Proceedings. AMIA Symposium
|February 5, 2002
PubMed
Summary

Data anonymization using cell suppression protects privacy but may not prevent individual inferences. Increased anonymity levels can degrade predictive model performance, impacting data utility.

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Training in Health Informatics in Brazil.

Yearbook of medical informatics·2016
Same author

Creating a Common Data Model for Comparative Effectiveness with the Observational Medical Outcomes Partnership.

Applied clinical informatics·2015
Same author

Health IT and clinical decision support systems: human factors and successful adoption.

Journal of the American Medical Informatics Association : JAMIA·2014
Same author

"Big data" and the electronic health record.

Yearbook of medical informatics·2014
Same author

Confluence of disciplines in health informatics: an international perspective.

Methods of information in medicine·2011
Same author

Feasibility evaluation of Smart Stretcher to improve patient safety during transfers.

Methods of information in medicine·2010

Area of Science:

  • Computer Science
  • Data Privacy
  • Statistical Analysis

Background:

  • Protecting individual data in disclosed databases is crucial.
  • Data anonymization techniques, such as table ambiguation via cell suppression, are employed to safeguard privacy.
  • The level of anonymization is determined by the required indistinguishability of individuals within the dataset.

Purpose of the Study:

  • To investigate the effectiveness of cell suppression in preventing inferences from anonymized data.
  • To analyze the trade-off between data anonymization levels and the preservation of descriptive characteristics.
  • To evaluate the impact of anonymization on the predictive performance of statistical models.

Main Methods:

  • Employed table ambiguation through selective cell suppression to achieve varying degrees of anonymization.

Related Experiment Videos

  • Assessed the ability of anonymized datasets to preserve descriptive data characteristics.
  • Quantified the potential for making inferences about individuals from anonymized data.
  • Measured the degradation of predictive model performance as a function of the anonymization level.
  • Main Results:

    • Anonymized datasets can retain descriptive data properties.
    • Cell suppression does not guarantee prevention of inferences, potentially compromising confidentiality.
    • A direct proportional relationship exists between the degree of anonymity and the degradation of predictive performance.
    • Demonstrated the effect of anonymization on a disease probability prediction model.

    Conclusions:

    • While data anonymization by cell suppression can obscure individuals, it may not fully prevent inferential attacks.
    • There is a quantifiable trade-off between achieving higher levels of anonymity and maintaining the predictive utility of data.
    • The findings highlight the need to carefully consider the desired balance between privacy protection and data usability in anonymization strategies.