Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Multiple Regression01:25

Multiple Regression

3.3K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.3K
Outliers and Influential Points01:08

Outliers and Influential Points

5.0K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
5.0K
Calculating and Interpreting the Linear Correlation Coefficient01:11

Calculating and Interpreting the Linear Correlation Coefficient

6.8K
The correlation coefficient, r, developed by Karl Pearson in the early 1900s, is numerical and provides a measure of strength and direction of the linear association between the independent variable, x, and the dependent variable, y. Hence, it is also known as the Pearson product-moment correlation coefficient. It can be calculated using the following equation:
6.8K
Cluster Sampling Method01:20

Cluster Sampling Method

13.3K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
13.3K
Residuals and Least-Squares Property01:11

Residuals and Least-Squares Property

8.2K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
8.2K
Regression Toward the Mean01:52

Regression Toward the Mean

6.6K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.6K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Determinants of high annual sickness absence in older workers: a prospective cohort study in England (Health and Employment After Fifty study).

BMJ open·2026
Same author

Navigating public health research in UK secondary schools: key challenges and opportunities identified by researchers.

BMC research notes·2026
Same author

The Association Between Maternal Adiposity and Breastfeeding Initiation and Duration: Evidence from the Southampton Women's Survey.

Maternal and child health journal·2026
Same author

Mixed-methods process evaluation of the EACH-B intervention in UK secondary schools: Delivery fidelity, stakeholder responses and contextual influences.

BMJ public health·2025
Same author

Process evaluation of a randomised controlled trial aimed at improving health behaviours and vitamin D status during pregnancy: Implementation of the SPRING trial.

PloS one·2025
Same author

Postnatal Growth Trajectories and Risk of Obstructive Sleep Apnea in Middle Age: A Cohort Study.

Pediatric pulmonology·2024

Related Experiment Video

Updated: Oct 29, 2025

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
12:27

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations

Published on: February 15, 2017

7.1K

Consequences of ignoring clustering in linear regression.

Georgia Ntani1,2, Hazel Inskip3,4, Clive Osmond3

  • 1Medical Research Council Lifecourse Epidemiology Unit, University of Southampton, Southampton, UK. gn@mrc.soton.ac.uk.

BMC Medical Research Methodology
|July 8, 2021
PubMed
Summary

Ignoring clustered data in regression analysis can lead to misleading conclusions. Ordinary least squares (OLS) regression is particularly prone to errors when outcome data are highly clustered, especially with continuous explanatory variables.

Keywords:
BiasClusteringComparisonConsequencesLinear regressionRandom intercept modelSimulation

More Related Videos

Assisted Selection of Biomarkers by Linear Discriminant Analysis Effect Size LEfSe in Microbiome Data
04:57

Assisted Selection of Biomarkers by Linear Discriminant Analysis Effect Size LEfSe in Microbiome Data

Published on: May 16, 2022

16.5K
Using Cholesky Decomposition to Explore Individual Differences in Longitudinal Relations between Reading Skills
06:52

Using Cholesky Decomposition to Explore Individual Differences in Longitudinal Relations between Reading Skills

Published on: September 17, 2019

6.5K

Related Experiment Videos

Last Updated: Oct 29, 2025

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
12:27

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations

Published on: February 15, 2017

7.1K
Assisted Selection of Biomarkers by Linear Discriminant Analysis Effect Size LEfSe in Microbiome Data
04:57

Assisted Selection of Biomarkers by Linear Discriminant Analysis Effect Size LEfSe in Microbiome Data

Published on: May 16, 2022

16.5K
Using Cholesky Decomposition to Explore Individual Differences in Longitudinal Relations between Reading Skills
06:52

Using Cholesky Decomposition to Explore Individual Differences in Longitudinal Relations between Reading Skills

Published on: September 17, 2019

6.5K

Area of Science:

  • Epidemiology
  • Clinical Research
  • Biostatistics

Background:

  • Clustering of observations is prevalent in epidemiological and clinical research.
  • Multilevel analysis is recommended for clustered data, but simpler methods are often used.
  • Failure to account for clustering can lead to erroneous conclusions in statistical inference.

Purpose of the Study:

  • To explore circumstances where ignoring clustering in linear regression leads to significant errors.
  • To compare the performance of random-intercept (RI) and ordinary least squares (OLS) models with clustered data.

Main Methods:

  • Simulated data based on the random-intercept model specification.
  • Varied scenarios of outcome clustering and explanatory variable types (continuous/binary).
  • Fitted RI and OLS models, comparing effect estimates, precision, confidence interval coverage, and Type I error rates.

Main Results:

  • Effect estimates were unbiased on average but showed greater deviation with increased outcome clustering.
  • OLS models showed larger deviations than RI models for continuous explanatory variables, especially with less clustering.
  • OLS precision was overestimated when explanatory variables varied more between clusters; confidence interval coverage and Type I error rates were problematic for OLS with continuous predictors.

Conclusions:

  • OLS regression can mislead statistical inference with clustered data.
  • Errors are most likely with continuous explanatory variables and highly clustered outcomes (ICC ≥ 0.01).