Related Experiment Video
Updated: Sep 16, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Robustly measuring multimorbidity using disparate linked datasets
Regina Prigge1, Kelly J Fleetwood2, Caroline A Jackson2
1Usher Institute, University of Edinburgh, Edinburgh, UK. regina.prigge@ed.ac.uk.
Background:
Measurement of multimorbidity, the co-occurrence of two or more conditions in the same individual, is highly variable which limits the consistency and reproducibility of research.
Methods:
Using data from 172,563 UK Biobank (UKB) participants and a cross-sectional approach, we examined how choice of data source affected estimated prevalence of 80 individual long-term conditions (LTCs) and multimorbidity. We developed code-list-based algorithms to determine the prevalence of 80 LTCs in (1) primary care records, (2) UKB baseline assessment, (3) hospital/cancer registry records, and (4) all three data sources together.
Results:
Using records from all three data sources, 146,811 (85.1%) participants have at least one and 109,609 (63.5%) have at least two LTCs at baseline. A median of 4.7% (IQR 1.0-16.6) of participants with a condition are identified by all three data sources. Agreement is highest for endocrine, nutritional and metabolic disorders, with a median of 32.9% (IQR 20.5-34.1) of individuals with a condition identified by all three data sources. Agreement is lowest for diseases of the genitourinary system and mental and behavioural disorders where perfect agreement varies from zero to 4.9% and zero to 12.3% across conditions, respectively. The low agreement between data sources is accompanied by high proportions of individuals with a condition identified only in primary care data (i.e. not in either of the other two sources), with a median of 59.3% (IQR 47.4-75.9) for diseases of the genitourinary system and 66.9% (IQR 42.8-79.2) for mental and behavioural disorders.
Conclusions:
Our study highlights the impact of the choice of which data source is used in research on individual LTCs and multimorbidity, and the importance of clearly justifying choices made.
More Related Videos
08:51Author Spotlight: Integrated Multi-Omics Analysis for Unveiling Multicellular Immune Signatures in Clinical Heart Attack Cohorts
Published on: September 20, 2024
07:35Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Related Concept Videos
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Statistical Methods for Analyzing Epidemiological Data
Genomics
Bias in Epidemiological Studies
Confounding in Epidemiological Studies
Principles of Disease Surveillance