Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Comparing the Survival Analysis of Two or More Groups01:20

Comparing the Survival Analysis of Two or More Groups

551
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
551
Multiple Regression01:25

Multiple Regression

3.7K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.7K
Residuals and Least-Squares Property01:11

Residuals and Least-Squares Property

9.1K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
9.1K
Regression Analysis01:11

Regression Analysis

8.0K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
8.0K
The Mantel-Cox Log-Rank Test01:19

The Mantel-Cox Log-Rank Test

993
The Mantel-Cox log-rank test is a widely used statistical method for comparing the survival distributions of two groups. It tests whether a statistically significant difference exists in survival times between the groups without assuming a specific distribution for the survival data, making it a non-parametric test. This flexibility makes the log-rank test particularly valuable in medical research and other fields where the timing of an event, such as death or disease recurrence, is of...
993
Parametric Survival Analysis: Weibull and Exponential Methods01:14

Parametric Survival Analysis: Weibull and Exponential Methods

1.0K
Parametric survival analysis models survival data by assuming a specific probability distribution for the time until an event occurs. The Weibull and exponential distributions are two of the most commonly used methods in this context, due to their versatility and relatively straightforward application.
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
1.0K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

A community case study of pedestrian environment features of neighborhood walkability and built environment-related 311 service requests surrounding public housing developments serving low-income Boston residents.

Frontiers in public health·2026
Same author

Multi-cancer early detection tests: national estimates of awareness and perceived value in the United States.

Cancer causes & control : CCC·2026
Same author

Facing the pandemic amidst mental health challenges: A longitudinal study of Black and Latino public housing residents.

Journal of affective disorders·2026
Same author

Syringe service program utilization, behavioral, and experiential factors associated with greater naloxone protection in a longitudinal cohort of people who use illicit opioids in New York city.

Harm reduction journal·2026
Same author

The Role of Harsh Discipline in Early Childhood Trajectories of Anxiety and Depressive Symptoms.

Academic pediatrics·2026
Same author

Police Pursuit Fatality Rates in the US and Directions for Future Research.

JAMA network open·2026

Related Experiment Video

Updated: Jan 14, 2026

Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
07:13

Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model

Published on: April 18, 2025

485

A comparison of methods for coding race in linear and logistic regression models.

Melody S Goodman1, Ariana Lopez2, Anarina L Murillo3

  • 1Department of Biostatistics, New York University School of Global Public Health, New York, NY, USA.

Annals of Epidemiology
|October 18, 2025
PubMed
Summary

Choosing how to code race in regression models significantly impacts findings on racial health inequities. The reference group is crucial for accurately measuring disparities in health outcomes, affecting public health research.

Keywords:
EncodingLinear regressionLogistic regressionRace variable

More Related Videos

Establishing a Competing Risk Regression Nomogram Model for Survival Data
04:57

Establishing a Competing Risk Regression Nomogram Model for Survival Data

Published on: October 23, 2020

10.7K
The Innovation Arena: A Method for Comparing Innovative Problem-Solving Across Groups
14:14

The Innovation Arena: A Method for Comparing Innovative Problem-Solving Across Groups

Published on: May 13, 2022

6.3K

Related Experiment Videos

Last Updated: Jan 14, 2026

Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
07:13

Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model

Published on: April 18, 2025

485
Establishing a Competing Risk Regression Nomogram Model for Survival Data
04:57

Establishing a Competing Risk Regression Nomogram Model for Survival Data

Published on: October 23, 2020

10.7K
The Innovation Arena: A Method for Comparing Innovative Problem-Solving Across Groups
14:14

The Innovation Arena: A Method for Comparing Innovative Problem-Solving Across Groups

Published on: May 13, 2022

6.3K

Area of Science:

  • Public Health
  • Biostatistics
  • Health Equity Research

Background:

  • Race is often controlled for in regression models as a confounder.
  • Traditional methods of coding race lack critical examination, potentially obscuring health inequities.

Purpose of the Study:

  • To compare various methods for coding race in regression analyses.
  • To highlight the impact of coding choices on identifying racial health disparities.

Main Methods:

  • Compared six race coding methods (dummy, simple effect, difference, deviation, analyst-defined) in linear and logistic regression.
  • Utilized data from 8097 participants in the 2020 New York City Community Health Survey.
  • Analyzed body mass index and diabetes status as outcome variables.

Main Results:

  • The selection of a race coding method influences the identification of racial health inequities.
  • The choice of the reference group is critical for accurately assessing racial disparities in health outcomes.

Conclusions:

  • Coding techniques in regression analysis have significant implications for research on racial health inequities.
  • Researchers must carefully consider the impact of race coding methods and reference group selection on study findings.