Related Experiment Video
Updated: Feb 15, 2026

In Vivo Protocol of Controlled Subconcussive Head Impacts for the Validation of Field Study Data
Published on: April 18, 2019
Valid and efficient subgroup analyses using nested case-control data
Bénédicte Delcoigne1, Nathalie C Støer2, Marie Reilly1
1Department of Medical Epidemiology and Biostatistics, Karolinska Institutet, Stockholm, Sweden.
Background:
It is not uncommon for investigators to conduct further analyses of subgroups, using data collected in a nested case-control design. Since the sampling of the participants is related to the outcome of interest, the data at hand are not a representative sample of the population, and subgroup analyses need to be carefully considered for their validity and interpretation.
Methods:
We performed simulation studies, generating cohorts within the proportional hazards model framework and with covariate coefficients chosen to mimic realistic data and more extreme situations. From the cohorts we sampled nested case-control data and analysed the effect of a binary exposure on a time-to-event outcome in subgroups defined by a covariate (an independent risk factor, a confounder or an effect modifier) and compared the estimates with the corresponding subcohort estimates. Cohort analyses were performed with Cox regression, and nested case-control samples or restricted subsamples were analysed with both conditional logistic regression and weighted Cox regression.
Results:
For all studied scenarios, the subgroup analyses provided unbiased estimates of the exposure coefficients, with conditional logistic regression being less efficient than the weighted Cox regression.
Conclusions:
For the study of a subpopulation, analysis of the corresponding subgroup of individuals sampled in a nested case-control design provides an unbiased estimate of the effect of exposure, regardless of whether the variable used to define the subgroup is a confounder, effect modifier or independent risk factor. Weighted Cox regression provides more efficient estimates than conditional logistic regression.
Related Concept Videos
Data Validation
Key parameters for method validation include:
Data Validation
Nursing assessment guides are generally based on holistic models rather than medical...
Reliability and Validity
Combinatorial Gene Control
The expression of more than 30,000 genes is controlled by approximately 2000-3000 transcription factors. This is possible because a single transcription factor can recognize more than one regulatory sequence. The specificity in gene...
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
How Data are Classified: Numerical Data
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...

