Related Experiment Video
Updated: Dec 9, 2025

14:06
Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
Published on: June 23, 2012
15.6K
Understanding between-cluster variation in prevalence and limits for how much variation is plausible
Mark D Chatfield1, Daniel M Farewell2
1Faculty of Medicine, The University of Queensland, Brisbane, Australia.
Statistical Methods in Medical Research
|September 10, 2020
Summary
Understanding between-cluster variation in clinical trials is crucial. New methods, including visualizing cluster prevalence distributions and using log odds scales, offer more intuitive and transportable measures than the intra-cluster correlation coefficient.
Area of Science:
- Biostatistics
- Clinical Trial Design
- Epidemiology
Background:
- Understanding between-cluster variation is essential for analyzing clustered binary data in clinical trials and observational studies.
- The intra-cluster correlation coefficient (ICC), commonly used in sample size and power calculations for cluster randomized trials, can be unintuitive.
- Low ICC values may mask substantial between-cluster differences, complicating interpretation.
Purpose of the Study:
- To propose more interpretable and transportable methods for quantifying between-cluster variation in clustered binary data.
- To provide practical guidance for clinical trialists in assessing and specifying between-cluster variation.
- To introduce bounds for key variation metrics based on maximum entropy theory.
Main Methods:
- Visualizing the distribution of true cluster prevalences, potentially assuming a beta distribution.
- Calculating the standard deviation of true cluster prevalences as a more interpretable alternative to ICC.
- Applying maximum entropy theory to derive rule-of-thumb bounds for ICC, standard deviation, coefficient of variation, and log odds scale metrics.
Main Results:
- The standard deviation of true cluster prevalences is more interpretable than ICC.
- Plausible ICC and standard deviation values are bounded by overall prevalence, its complement, and one-third.
- Metrics on the log odds scale demonstrated greater transportability across studies with varying prevalences compared to ICC and coefficient of variation.
Conclusions:
- Visualizing cluster prevalence distributions and using log odds scale metrics improve the understanding of between-cluster variation.
- The proposed bounds offer practical guidance for avoiding implausibly high ICC values in sample size calculations.
- This work enhances the interpretability and transportability of between-cluster variation measures in clustered binary data analysis.
Related Concept Videos
Variability: Analysis
349
Measures of variability are statistical metrics that reveal the dispersion pattern within a dataset. They are pivotal in biostatistics, providing insights into the heterogeneity within health and biological data. Variability signifies the degree to which data points diverge from one another, helping researchers understand the potential range of values and associated uncertainty within the data.
The range is a simple measure of variability, indicating the difference between the highest and...
The range is a simple measure of variability, indicating the difference between the highest and...
349
What is Variation?
16.7K
Apart from the measures of central tendency, distribution, outliers, and the changing characteristics of data with time, an important characteristic of any data set is its variation or spread. In some data sets, the data values are concentrated closely near the mean; in others, the data values are more widely spread out from the mean.
The range, standard deviation, standard error, and variance are the different measures of variation.
Range: The range is the difference between its maximum and...
The range, standard deviation, standard error, and variance are the different measures of variation.
Range: The range is the difference between its maximum and...
16.7K
Uncertainty: Confidence Intervals
9.2K
The confidence interval is the range of values around the mean that contains the true mean. It is expressed as a probability percentage. The interpretation of a 95% confidence interval, for instance, is that the statistician is 95% confident that the true mean falls within the interval. The upper and lower limits of this range are known as confidence limits. The confidence limits for the true mean are estimated from the sample's mean, the standard deviation, and the statistical factor...
9.2K
Unusual Results
3.7K
Unusual results are those that have a very low chance of occurring. Unusual results can be identified using probabilities and the range rule of thumb. In problems involving probability, unusual results can be observed in 2 instances – an unusually high number of successes or an unusually low number of successes.
According to the range rule of thumb, any value above or below two standard deviations, 2σ from the mean, μ is considered unusual.
Maximum unusual value =...
According to the range rule of thumb, any value above or below two standard deviations, 2σ from the mean, μ is considered unusual.
Maximum unusual value =...
3.7K
Confidence Intervals
9.7K
An unbiased point estimate is often insufficient to predict a population estimate, such as population mean or population proportion. In this scenario, a confidence interval is used. A confidence interval is an estimate similar to a sample proportion. However, unlike the point estimate which is a single value, the confidence interval contains a range of values. These values have lower and upper limits, known as confidence limits, and can be designated as L1 and L2, respectively.
A...
A...
9.7K
Interpretation of Confidence Intervals
8.9K
A confidence interval is a better estimate of the population than a point estimate, as it uses a range of values from a sample instead of a single value.
Confidence intervals have confidence coefficients that are crucial for their interpretation. The most common confidence coefficients are 0.90, 0.95, and 0.99, which can be written as percentages–90%, 95%, and 99%, respectively.
Suppose a person calculates a confidence interval with a confidence coefficient of 0.95. In that case, they can...
Confidence intervals have confidence coefficients that are crucial for their interpretation. The most common confidence coefficients are 0.90, 0.95, and 0.99, which can be written as percentages–90%, 95%, and 99%, respectively.
Suppose a person calculates a confidence interval with a confidence coefficient of 0.95. In that case, they can...
8.9K

