Methods for observed-cluster inference when cluster size is informative: a review and clarifications
Shaun R Seaman1, Menelaos Pavlou, Andrew J Copas
1MRC Biostatistics Unit, Cambridge CB2 0SR, UK.
Biometrics
|February 1, 2014
Summary
Epidemiological studies with missing data require careful analysis. Existing methods for clustered data may provide estimates for complete clusters, not observed clusters, potentially misinterpreting results.
Area of Science:
- Epidemiology
- Biostatistics
- Statistical Modeling
Background:
- Clustered data are common in epidemiology, often with missing outcome data.
- Missing outcomes can arise from non-existence (e.g., death) or other reasons.
- The distribution of outcomes in complete clusters may differ from observed clusters.
Purpose of the Study:
- To evaluate statistical methods for clustered data with missing outcomes.
- To determine if existing methods provide inference for observed or complete clusters.
- To identify conditions where shared random-effects models yield valid observed-cluster inference.
Main Methods:
- Review of weighted and doubly weighted generalized estimating equations.
- Analysis of shared random-effects models for clustered data.
- Investigation of conditions for observed-cluster inference.
Main Results:
- Proposed methods for observed-cluster inference may actually provide inference for complete clusters.
- This holds true even when observed clusters are internally complete.
- Conditions are identified for shared random-effects models to accurately reflect observed data.
Conclusions:
- Misinterpretation of estimates from shared random-effects models is a significant risk.
- Careful consideration of complete vs. observed cluster inference is crucial.
- Psoriatic arthritis data illustrate the potential for misleading results.
Related Concept Videos
Cluster Sampling Method
11.0K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.0K
Sampling Plans
1.5K
Sampling is a crucial step in analytical chemistry, allowing researchers to collect representative data from a large population. Common sampling methods include random, judgmental, systematic, stratified, and cluster sampling.
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
1.5K
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
729
Statistical inference techniques, paramount in hypothesis testing, differentiate into two broad categories: parametric and nonparametric statistics.
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
729
Interpretation of Confidence Intervals
8.8K
A confidence interval is a better estimate of the population than a point estimate, as it uses a range of values from a sample instead of a single value.
Confidence intervals have confidence coefficients that are crucial for their interpretation. The most common confidence coefficients are 0.90, 0.95, and 0.99, which can be written as percentages–90%, 95%, and 99%, respectively.
Suppose a person calculates a confidence interval with a confidence coefficient of 0.95. In that case, they can...
Confidence intervals have confidence coefficients that are crucial for their interpretation. The most common confidence coefficients are 0.90, 0.95, and 0.99, which can be written as percentages–90%, 95%, and 99%, respectively.
Suppose a person calculates a confidence interval with a confidence coefficient of 0.95. In that case, they can...
8.8K
Statistical Analysis: Overview
14.6K
When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
14.6K
One-Way ANOVA: Equal Sample Sizes
3.2K
One-Way ANOVA can be performed on three or more samples with equal or unequal sample sizes. When one-way ANOVA is performed on two datasets with samples of equal sizes, it can be easily observed that the computed F statistic is highly sensitive to the sample mean.
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
3.2K


