Related Experiment Video
Updated: Jun 28, 2025

Using the Race Model Inequality to Quantify Behavioral Multisensory Integration Effects
Published on: May 10, 2019
Categorization of continuous covariates and complex regression models-friends or foes in intersectionality research
Adrian Richter1, Sabina Ulbricht1, Sarah Brockhaus2
1Department of Prevention Research and Social Medicine, Institute for Community Medicine, University Medicine Greifswald, Greifswald, Germany.
Objectives:
To reduce health inequities, it is important to identify intersections in characteristics of individuals subject to privilege or disadvantage. Different proposals for that have recently been published. One approach (1) considers models specified with first- and all second-order effects and another (2) the stratification based on multiple covariates; both categorize continuous covariates. A simulation study was conducted in order to review both methods with regard to identification of intersections showing true differences, rate of false-positive results, and generalizability to independent data compared to an established approach (3) of backward variable elimination according to Bayesian information criterion (BE-BIC) combined with splines.
Study Design And Setting:
R software has been used to simulate the covariates age, sex, body mass index, education, and diabetes to examine their association with a continuous frailty score for osteoporosis using multiple linear regression. In setting 1, none of the covariates was associated with the frailty score, that is, only noise is present in the data. In setting 2, the covariates age, sex, and their interaction were associated with the frailty score, such that only females above 55 years formed an intersection associated with an increased frailty score. All approaches were compared under varying sample sizes (N = 200-3000) and signal-to-noise ratios (SNRs, 0.5-4) in 1000 replications. For model evaluation, bootstrap resampling was used. The models were fitted in internal learning data and then used to predict outcomes in the internal validation data. The mean squared error (MSE) was used for comparison and the frequency of false-positive findings calculated.
Results:
In setting 1, approaches 1 and 2 generated spurious effects in more than 90% of simulations across all sample sizes. In a smaller sample size, approach 3 (BE-BIC) selected 36.5% of the correct model, in larger sample size in 89.8% and always had a lower number of spurious effects. MSE in independent data was generally higher for approaches 1 and 2 when compared to 3. In setting 2, approach 1 selected most frequently the correct interaction but frequently showed spurious effects (>75%). Across all sample sizes and SNR, approach 3 generated least often spurious results and had lowest MSE in independent data.
Conclusion:
Categorization of continuous covariates is detrimental to studies on intersectionality. Due to high and unrestricted model complexity, such approaches are prone to spurious effects and often lack interpretability. Approach 3 (BE-BIC) is considerably more robust against spurious findings, showed better generalizability to independent data, and can be used with most statistical software. For intersectionality research, we consider it most important to describe relevant differences between intersections and to avoid nonreproducible and spurious findings.
More Related Videos
Related Concept Videos
Multicompartment Models: Overview
These models offer a more comprehensive representation of drug behavior in the body than one-compartment models. They accommodate the complexity of drug distribution,...
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Mechanistic Models: Compartment Models in Individual and Population Analysis
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Confounding in Epidemiological Studies
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:

