Related Experiment Video
Updated: May 9, 2026

Using the Race Model Inequality to Quantify Behavioral Multisensory Integration Effects
Published on: May 10, 2019
Geographic Variation in Missing Race and Ethnicity Data in Minimum Data Set 3.0
Jennifer Tjia1, Francesca L Troiani1, Anna Wyndham2
1Department of Population and Quantitative Health Sciences, Division of Epidemiology, Worcester, MA.
Background:
Race and ethnicity measures in administrative data can vary geographically. The extent of this challenge in US nursing homes is not well described.
Objectives:
To describe geographic variation in missing race and ethnicity data in the Minimum Data Set (MDS) 3.0 and Medicare claims, and to compare discrepancies across data sources.
Research Design:
Cross-sectional study.
Subjects:
Medicare beneficiaries with MDS 3.0 records between 2014 and 2018. The Medicare Beneficiary Summary File provided demographic information.
Measures:
Missingness of MDS race and ethnicity data by state, and misclassification of Medicare race and ethnicity enrollment database (EDB) and Research Triangle Institute (RTI) variables compared with MDS. We calculate the sensitivity, specificity, and positive predictive value of the EDB and RTI variables relative to the MDS.
Results:
Among 18.1 million nursing home residents pooled across 2014-2018, geographic variation in missing race and ethnicity in the MDS 3.0 ranged from 1.2% to 14.7%. Compared with MDS, misclassification of residents classified as Hispanic in MDS ranged from 48.1% to 89.2% for EDB and 0.5% to 44.8% for RTI. Misclassification of residents classified as Asian American/Pacific Islander in MDS ranged from 29.4% to 77.2% for EDB and 12.7% to 65.4% for RTI. Misclassification of residents classified as Black ranged from 0% to 14.2% for EDB and 0% to 16.2% for RTI. Overall, the RTI variables provided better sensitivity and specificity of race and ethnicity than the EDB.
Conclusion:
Missing race and ethnicity data in the MDS varies geographically, as do discrepancies between MDS and EDB and RTI variables. Thoughtful consideration of these issues is recommended when handling missing MDS race and ethnicity data.
Insights
Geographic variation exists in missing race and ethnicity data within US nursing homes. Discrepancies between Minimum Data Set (MDS) and Medicare data sources highlight data quality challenges.
Area of Science:
- Gerontology
- Health Services Research
- Data Science
Background:
- Administrative data on race and ethnicity in US nursing homes exhibit geographic variability.
- The extent of missing race and ethnicity data in nursing home administrative records is not well-characterized.
Purpose of the Study:
- To analyze geographic variations in missing race and ethnicity data within the Minimum Data Set (MDS) 3.0.
- To compare discrepancies in race and ethnicity data between MDS 3.0 and Medicare claims.
Main Methods:
- A cross-sectional study analyzed Medicare beneficiaries with MDS 3.0 records from 2014-2018.
- Missingness of race and ethnicity data in MDS was assessed by state.
- Misclassification rates of race and ethnicity variables from Medicare enrollment database (EDB) and Research Triangle Institute (RTI) were compared against MDS data.
Main Results:
- Significant geographic variation in missing MDS race and ethnicity data was observed, ranging from 1.2% to 14.7%.
- Misclassification rates varied substantially between MDS and Medicare data sources (EDB and RTI) for Hispanic, Asian American/Pacific Islander, and Black populations.
- RTI variables demonstrated superior sensitivity and specificity compared to EDB when contrasted with MDS data.
Conclusions:
- Geographic disparities in missing race and ethnicity data are evident in US nursing homes' MDS.
- Discrepancies between MDS and Medicare data sources (EDB, RTI) also vary geographically.
- Careful consideration of these data quality issues is crucial when utilizing MDS race and ethnicity information.
Related Concept Videos
What is Variation?
The range, standard deviation, standard error, and variance are the different measures of variation.
Range: The range is the difference between its maximum and...
Midrange
Simply put, the midrange is half of the data set’s range. Similar to the mean, the midrange is sensitive to the extreme values and hence the prospective outliers. However, unlike the mean, the midrange is not sensitive to all the values of the data set that lie in the middle. Thus, it is prone to outliers and...
Variation: Normal Distribution, Range, and Standard Deviation
Trimmed Mean
Although certain measures of central tendency are not sensitive to outliers, there are alternative versions of the mean that get around the...
One-Way ANOVA: Equal Sample Sizes
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
