Using machine learning to predict future foster care admission
Ari Ne'eman1, Alex Brooks2, Kellie Hans-Green2
1Department of Health Policy and Management, Harvard T.H. Chan School of Public Health, Boston, MA 02115, United States.
Introduction:
Foster care admissions are highly traumatic for children and their families, often causing serious adverse outcomes. We seek to assess the viability of machine learning methods to identify children at risk of future foster care admission to facilitate diversion.
Methods:
We use claims data for children enrolled in a Medicaid health plan in Ohio as well as for linked adults, along with data on individual and geographic social determinants of health (SDOH) factors. We test the performance of a gradient-boosted tree machine learning algorithm as compared to logistic regression. Of the children, 85% have SDOH data available.
Results:
Using a gradient-boosted tree machine learning algorithm, we built a model that identifies 2408 children (1.32%) as at risk of foster care admission in a sample of 181 841, of whom 1599 entered foster care within 1 year, resulting in a positive predictive value (PPV) of 66.4% (F 1 = 55.5%, specificity = 99.5%, sensitivity = 47.67%), outperforming logistic regression. Accuracy was substantially better when using SDOH data (PPV of 84.72% with SDOH data compared to 27.44% without).
Conclusions:
These results highlight the importance of SDOH factors in predicting foster care admission. They also point to the potential of machine learning for facilitating early intervention to prevent foster care admissions.
Related Concept Videos
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
End Point Prediction: Gran Plot
For potentiometric titration, the Gran plot is created by plotting...
Predicting Reaction Outcomes
Steps in Outbreak Investigation
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
