Related Experiment Video
Updated: Mar 17, 2026

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Privacy preserving data anonymization of spontaneous ADE reporting system dataset
Wen-Yang Lin1, Duen-Chuan Yang2, Jie-Teng Wang2
1Department of Computer Science and Information Engineering, National University of Kaohsiung, Nanzih District, Kaohsiung, 811, Taiwan, R.O.C. wylin@nuk.edu.tw.
This study introduces a novel privacy model, MS(k, θ(*))-bounding, for anonymizing spontaneous reporting system (SRS) data. The method effectively protects sensitive health information while preserving data utility for adverse drug reaction (ADR) detection.
Area of Science:
- Pharmacovigilance and Drug Safety
- Data Privacy and Anonymization
- Health Informatics
Background:
- Spontaneous Reporting Systems (SRSs) are crucial for long-term drug safety surveillance.
- Publishing SRS data raises privacy concerns due to sensitive personal health information.
- Existing privacy-preserving data publishing (PPDP) methods are inadequate for SRS data characteristics like rare events and multi-valued attributes.
Purpose of the Study:
- To propose a new privacy model, MS(k, θ(*))-bounding, specifically designed for spontaneous adverse drug event (ADE) reporting data.
- To develop an anonymization algorithm that addresses the unique challenges of SRS datasets.
- To evaluate the effectiveness of the proposed model in balancing privacy protection and data utility.
Main Methods:
- Developed the MS(k, θ(*))-bounding privacy model with flexible privacy thresholds (θ(*)) for varying sensitive values.
- Proposed a greedy-based clustering anonymization algorithm to minimize privacy risk and maintain data utility.
- Conducted empirical studies using the FAERS dataset (2004Q1-2011Q4) and compared with k-anonymity, (X, Y)-anonymity, Multi-sensitive l-diversity, and (α, k)-anonymity.
- Evaluated methods using Danger Ratio (DR) and Information Loss (IL) across uniform, level-wise, and frequency-based threshold settings.
Main Results:
- The MS(k, θ(*))-bounding model effectively prevented sensitive value disclosure (near-zero DRs) across all threshold settings.
- Non-uniform threshold settings (level-wise, frequency-based) demonstrated superior data utility and minimal privacy risk compared to other models.
- Anonymized data showed minimal impact on the strength of discovered adverse drug reaction (ADR) signals (e.g., PRR, ROR).
Conclusions:
- A novel privacy model and anonymization algorithm were developed for protecting SRS data with unique characteristics.
- The proposed method effectively addresses privacy concerns in SRS data without compromising data utility for ADR signal detection.
- Empirical evaluation on the FAERS dataset validates the method's efficacy in real-world scenarios.
Related Concept Videos
Pharmacovigilance
This process, termed pharmacovigilance, aims to detect, evaluate, and minimize harmful effects related to medication use. The data collection for pharmacovigilance depends on spontaneous reporting systems, where healthcare professionals or patients voluntarily report suspected ADRs.
In some cases, there...
Pharmaceutical Poisoning: Potential Scenarios
Censoring Survival Data
Analysis of Population Pharmacokinetic Data
Dosage Regimens: Partial Pharmacokinetic Parameters
Model-Independent Approaches for Pharmacokinetic Data: Noncompartmental Analysis
One important characteristic of noncompartmental analyses is that drug exposure increases proportionally with increasing doses. This...
