Related Experiment Video
Updated: Jan 1, 2026

06:55
Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
15.0K
Fair regression for health care spending
1PhD Program in Health Policy, Harvard University, Cambridge, Massachusetts.
Biometrics
|December 21, 2019
Summary
New fair regression methods improve health insurance risk adjustment for undercompensated groups, enhancing fairness by 98% with minimal impact on overall accuracy.
Area of Science:
- Health economics
- Computer science
- Statistics
Background:
- Health insurance risk adjustment formulas predict healthcare spending to ensure fair payments.
- Current formulas underpredict spending for certain groups, leading to undercompensation and potential enrollment discrimination.
- This impacts healthcare access for vulnerable populations.
Purpose of the Study:
- To develop improved risk adjustment methods addressing fairness for undercompensated groups.
- To integrate fairness considerations directly into regression models.
- To propose a novel fairness metric for evaluating risk adjustment formulas.
Main Methods:
- Developed new fair regression methods for continuous outcomes by incorporating fairness into the objective function.
- Utilized concepts from statistics, computer science, and health economics.
- Applied methods to IBM MarketScan Research Databases and conducted simulation studies.
Main Results:
- New fair regression methods demonstrated significant improvements in group fairness, achieving up to 98% enhancement.
- Overall model fit experienced only a minor reduction, approximately 4%.
- Proposed a novel fairness metric, emphasizing the need for a suite of metrics.
Conclusions:
- The developed fair regression methods offer substantial improvements in risk adjustment fairness for undercompensated groups.
- These methods can mitigate incentives for insurers to avoid enrolling specific populations.
- This research contributes to more equitable health insurance markets and improved healthcare access.
Related Concept Videos
Regression Toward the Mean
6.8K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.8K
Residuals and Least-Squares Property
8.8K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
8.8K
Statistical Methods for Analyzing Epidemiological Data
847
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
847
Multiple Regression
3.7K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.7K
Regression Analysis
7.7K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
7.7K
Microsoft Excel: Regression Analysis
1.4K
Regression analysis in Microsoft Excel is a powerful statistical method for examining the relationship between a dependent variable and one or more independent variables. It's used extensively in fields such as economics, biology, and business to predict outcomes, understand relationships, and make data-driven decisions. The most common type is linear regression, which attempts to fit a straight line through the data points to model the relationship between variables.
To perform regression...
To perform regression...
1.4K

