Related Experiment Video
Updated: May 27, 2026

Inverse Probability of Treatment Weighting (Propensity Score) using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
A bootstrapping algorithm to improve cohort identification using structured data
Sasikiran Kandula1, Qing Zeng-Treitler1, Lingji Chen2
1Department of Biomedical Informatics, University of Utah, Salt Lake City, UT, United States.
Abstract:
Cohort identification is an important step in conducting clinical research studies. Use of ICD-9 codes to identify disease cohorts is a common approach that can yield satisfactory results in certain conditions; however, for many use-cases more accurate methods are required. In this study, we propose a bootstrapping method that supplements ICD-9 codes with lab results, medications, etc. to build classification models that can be used to identify cohorts more accurately. The proposed method does not require prior information about the true class of the patients. We used the method to identify Diabetes Mellitus (DM) and Hyperlipidemia (HL) patient cohorts from a database of 800 thousand patients. Evaluation results show that the method identified 11,000 patients who did not have DM related ICD-9 codes as positive for DM and 52,000 patients without HL codes as positive for HL. A review of 400 patient charts (200 patients for each condition) by two clinicians shows that in both the conditions studied, the labeling assigned by the proposed approach is more consistent with that of the clinicians compared to labeling through ICD-9 codes. The method is reasonably automated and, we believe, holds potential for inexpensive, more accurate cohort identification.
Related Concept Videos
Bootstrapping
Stratified Sampling Method
To choose a stratified sample, divide the population into groups called strata and then take a...
Systematic Sampling Method
Systematic sampling is one of the simplest methods...
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Statistical Methods for Analyzing Epidemiological Data
Longitudinal Studies