Quantile Regression in Epidemiology: Capturing Heterogeneity Beyond the Mean
1Department of Fisheries and Aquaculture, School of Agricultural Sciences, University of Patras, 30200 Messolonghi, Greece.
Abstract:
Ordinary linear regression is the most common approach for modeling relationships between continuous outcomes and explanatory variables in epidemiological research. However, this method relies on restrictive assumptions-normality, homoscedasticity, and linearity-that are often violated in real-world biomedical data. When these assumptions fail, mean-based estimates may obscure important heterogeneity across the outcome distribution. This study aims to illustrate the methodological and interpretive advantages of quantile regression over ordinary regression in the analysis of epidemiological data. Secondary data were derived from a cross-sectional study of 1415 healthy Greek adults aged 25-82 years. Body mass index (BMI) served as the outcome variable, while sex, age, physical activity, dieting status, and daily energy intake were considered predictors. Both ordinary and quantile regression models were applied to estimate associations between BMI and its determinants across the 25th, 50th, 75th, and 90th quantiles. Ordinary regression identified positive associations of BMI with age and energy intake and a negative association with physical activity. Quantile regression revealed that these relationships were not constant across the BMI distribution. The inverse association with physical activity intensified at higher quantiles, and the gender effect reversed direction at the upper tail, suggesting heterogeneity was not captured by mean-based models. Quantile regression provides a distribution-sensitive alternative to ordinary regression, offering insight into covariate effects across different points of the outcome distribution and serving as both a robust analytical tool and an educational framework for applied epidemiological research.
Related Concept Videos
Regression Toward the Mean
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Correlation and Regression
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Introduction to Epidemiology
Causality in Epidemiology


