Related Experiment Video
Updated: Mar 7, 2026

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Utilising identifier error variation in linkage of large administrative data sources
Katie Harron1, Gareth Hagger-Johnson2, Ruth Gilbert3
1London School of Hygiene and Tropical Medicine, 15-17 Tavistock Place, London, WC1 H 9SH, UK. Katie.harron@lshtm.ac.uk.
We found significant variation in patient data error rates across hospitals and demographics. Adjusting data linkage weights based on these variations can reduce bias in health research.
Area of Science:
- Health Informatics
- Data Science
- Biostatistics
Background:
- Probabilistic methods for linking administrative data rely on independent identifiers.
- Data quality variations can cause identifier errors, violating independence assumptions and introducing bias.
- This study addresses identifier error variation in Hospital Episode Statistics (HES).
Purpose of the Study:
- To measure identifier error rate variations in HES data.
- To develop and evaluate new methods for match weight calculation incorporating identifier error information.
- To reduce selection bias in linked administrative datasets.
Main Methods:
- Linked 30,000 HES records with Personal Demographic Service (PDS) data.
- Calculated error rates for sex, date of birth, and postcode.
- Used multi-level logistic regression to assess individual and organizational factors influencing errors.
- Derived attribute- and organization-specific match weights.
Main Results:
- Found significant variations in error rates by age, ethnicity, and sex (p < 0.0005).
- Postcode errors were prevalent (53%), while sex and DOB errors were low (0.11%).
- Simulation showed attribute- and organization-specific weights reduced bias compared to traditional methods.
Conclusions:
- Empirical evidence confirms significant identifier error variation in administrative data.
- A novel method for match weight derivation incorporating data variation is proposed.
- Accounting for individual-level characteristics in linkage weights can mitigate bias.
Related Concept Videos
Systematic Error: Methodological and Sampling Errors
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Documentation of Nursing Diagnosis
In some settings, data-driven computerized decision support systems are in place, allowing for more accurate nursing diagnoses. The database within one of these systems includes diagnostic labels defining characteristics, activities, and indicators for nursing. A nurse enters...
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Statistical Methods for Analyzing Epidemiological Data
Types of Records II: Educational and Administrative Records
Confounding in Epidemiological Studies

