Related Experiment Video
Updated: Jun 14, 2026

14:27
Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data
Published on: June 26, 2013
A method for managing re-identification risk from small geographic areas in Canada
Khaled El Emam1, Ann Brown, Philip AbdelMalik
1Children's Hospital of Eastern Ontario Research Institute, 401 Smyth Road, Ottawa, Ontario K1J 8L1, Canada. kelemam@uottawa.ca
BMC Medical Informatics and Decision Making
|April 6, 2010
Summary
New models help manage re-identification risks in health data by defining "small geographic areas" using uniqueness thresholds. Data custodians can select 0%, 5%, or 20% uniqueness levels based on data sensitivity and recipient controls.
Area of Science:
- Health data privacy
- Statistical disclosure control
- Geographic information systems
Background:
- Common practice involves suppressing or aggregating data from small geographic areas to protect privacy.
- The uniqueness criterion suggests an area is "small" if unique individuals (based on quasi-identifiers) approach a certain threshold.
- Existing thresholds for health data uniqueness range from 0% to 20%.
Purpose of the Study:
- To develop and validate predictive models for determining when geographic areas exceed specific uniqueness thresholds (5% and 20%).
- To provide data custodians with guidance on selecting appropriate uniqueness thresholds based on data characteristics and recipient capabilities.
- To manage the risk of re-identification in health datasets originating from small geographic areas.
Main Methods:
- Estimated uniqueness for urban Forward Sortation Areas (FSAs) using Canadian census data (20% population).
- Constructed logistic regression models to predict uniqueness above 5% and 20% thresholds.
- Validated model accuracy using 10-fold cross-validation, with population size and quasi-identifier equivalence classes as predictors.
Main Results:
- Models demonstrated high prediction accuracy (specificity > 0.9) and significant parameters.
- Sensitivity was 0.87 for the 5% threshold and 0.74 for the 20% threshold.
- Higher thresholds identified fewer records as "small," reducing disclosure control actions; guidance provided for threshold selection.
Conclusions:
- Developed models effectively manage re-identification risks associated with small geographic areas.
- Offers flexibility for data custodians to define "small geographic areas" based on data type and recipient.
- Enables tailored disclosure control strategies balancing privacy protection and data utility.
Related Concept Videos
Sampling Plans
Sampling is a crucial step in analytical chemistry, allowing researchers to collect representative data from a large population. Common sampling methods include random, judgmental, systematic, stratified, and cluster sampling.
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Selected Data About Geographic Locations
Geographic Information Systems (GIS) rely on two core types of data: spatial data and attribute data.Spatial DataSpatial data defines the physical location of features within a coordinate system, typically expressed in terms of latitude and longitude. It provides precise positioning for elements like roads, rivers, or buildings.Attribute DataAttribute data complements spatial data by adding descriptive information about these features. For example, a road's spatial data includes its start and...
Conservation of Small Populations
Small population sizes put a species at extreme risk of extinction due to a lack of variation, and a consequent decrease in adaptability. This weakens the chances of survival under pressures such as climate change, competition from other species, or new diseases. Large populations are more likely to survive pressures such as these, as such populations are more likely to harbor individuals that have genetic variants that are adaptive under new stresses. Small populations are much less likely to...
Levels of Use of a GIS
Geographic Information Systems (GIS) operate across three levels of application, each representing an increasing degree of complexity: data management, analysis, and prediction. These levels reflect the expanding functionality and versatility of GIS technology in handling spatial data for diverse purposes.Data ManagementAt its foundational level, GIS serves as a tool for data management, enabling the input, storage, retrieval, and organization of spatial data. This level is often employed in...
