Related Experiment Video
Updated: Aug 3, 2025

Using the Race Model Inequality to Quantify Behavioral Multisensory Integration Effects
Published on: May 10, 2019
Methods for retrospectively improving race/ethnicity data quality: a scoping review
Matthew K Chin1, Lan N Đoàn1, Rienna G Russo1
1Section for Health Equity, Department of Population Health, NYU Grossman School of Medicine, New York, NY 10016, United States.
Improving race and ethnicity data quality is crucial for identifying health disparities and informing policy. This review identified six methods for retrospectively improving classification in secondary data sets.
Area of Science:
- Health Informatics
- Data Science
- Public Health
Background:
- Accurate race and ethnicity data are vital for identifying health disparities and informing equitable healthcare policy.
- Secondary data sets often have incomplete or inaccurate race/ethnicity classifications, potentially misrepresenting underserved populations.
- Retrospective methods to improve race/ethnicity data quality are essential for robust health research.
Purpose of the Study:
- To conduct a scoping review of methods used to retrospectively improve race and ethnicity classification in secondary data sets.
- To identify and categorize different approaches for enhancing demographic data quality.
- To analyze the characteristics, strengths, and limitations of existing methods.
Main Methods:
- Systematic search of MEDLINE, Embase, and Web of Science Core Collection databases.
- Screening of 2,441 abstracts and full-text review of 453 articles, including 120 in the final analysis.
- Narrative analysis of study characteristics and categorization of identified methods.
Main Results:
- Six primary method types were identified: name lists (23%), name algorithms (46%), machine learning (12%), expert review (8%), data linkage (8%), and other (5%).
- The majority of studies focused on classifying Asian (47%) and White (43%) populations.
- Validation evaluations were present in 72% of the included articles.
Conclusions:
- Name algorithms are the most common method for improving race/ethnicity data, but innovative approaches are needed for better subgroup identification.
- Strengths, limitations, and potential harms of various methods require careful consideration.
- Accurate, disaggregated race/ethnicity data are critical for effective policymaking and healthcare interventions.
More Related Videos
06:05The Participant-Reported Implementation Update and Score PRIUS: A Novel Method for Capturing Implementation-Related Data Over Time
Published on: February 19, 2021
06:55Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020