Related Experiment Video
Updated: Jul 4, 2026

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
Published on: February 15, 2017
A comparison of regression approaches for analyzing clustered data
Manisha Desai1, Melissa D Begg
1Department of Biostatistics, Columbia University, 722 West 168th St, R646, New York, NY 10032, USA. manisha.desai@columbia.edu
Choosing the right statistical model for clustered data is crucial. Different analytical approaches significantly impact the interpretation of results, especially when examining factors like head circumference and intelligence (IQ).
Area of Science:
- Biostatistics
- Epidemiology
- Developmental Psychology
Background:
- Clustered data, such as sibling or family data, requires specialized statistical methods for accurate analysis.
- Model selection can influence the interpretation of associations between variables when data are not independent.
Purpose of the Study:
- To evaluate the impact of different analytical approaches on interpreting results from clustered data.
- To compare three distinct methods for analyzing clustered data, focusing on model specification and covariate handling.
Main Methods:
- Three analytical approaches were employed: two random intercept models with varying covariate specifications and one standard analysis of paired differences.
- The methods were applied to data from the National Collaborative Perinatal Project.
- The study examined the association between head circumference at birth and intelligence (IQ) at age 7.
Main Results:
- An approach ignoring within- and between-family effects estimated an overall IQ increase of 1.1 points per cm increase in head circumference.
- Two other approaches, accounting for within-family effects, found comparable results: 0.6 points (95% CI = 0.4, 0.9) and 0.69 points (95% CI = 0.4, 1.0).
Conclusions:
- Appropriate analytical methods are essential for correctly interpreting clustered data.
- Careful covariate specification in regression modeling is critical for distinguishing between within- and between-cluster effects.
- The choice of statistical method should align with research interest in cluster-level versus item-level effects.
Related Concept Videos
Comparing the Survival Analysis of Two or More Groups
Statistical Analysis: Overview
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
Statistical Methods to Analyze Parametric Data: ANOVA
One-way ANOVA is applied when a single independent variable or factor is scrutinized. It compares the...
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
