A comparison of methods for coding race in linear and logistic regression models
Melody S Goodman1, Ariana Lopez2, Anarina L Murillo3
1Department of Biostatistics, New York University School of Global Public Health, New York, NY, USA.
Abstract:
In many public health and clinical research studies that use regression models for analyses, race is often considered a confounder and "controlled" for in the regression model with simple indicators for race and non-Hispanic White as the reference group, without much introspection from the data analyst. From a health equity perspective, multiple issues exist with this approach. We examine and compare several methods for coding race in linear and logistic regression models. We compare several coding methods using a sample of 8097 participants (≥18 years old) from the 2020 New York City Community Health Survey. To illustrate the importance of coding methods for race, we conducted regression analyses to compare the results from six coding approaches: dummy, simple effect, difference (forward and backward), deviation, and analyst-defined coding. Body mass index measured continuously and diabetes status measured dichotomously were the outcome variables in the linear and logistic regression models. Results showed that selecting a coding method has implications for identifying racial health inequities. The reference group selection is critical to measuring racial inequities in health outcomes. This study emphasizes the need to consider the impact of coding techniques on research study design, particularly when racial health inequities are the research focus.
Related Concept Videos
Comparing the Survival Analysis of Two or More Groups
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
The Mantel-Cox Log-Rank Test
Parametric Survival Analysis: Weibull and Exponential Methods
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...


