使用机器学习方法预测美国大城市的低出生体重
1Department of Family Medicine and Siaal Research Center for Family Practice and Primary Care, The Haim Doron Division of Community Health, Faculty of Health Sciences, Ben-Gurion University of the Negev, Beer-Sheva 84161, Israel.
概括
机器学习准确地预测了美国主要城市的低出生体重率. 关键预测因素包括克拉米迪亚的发病率,种族隔离,产前护理的准入,单亲家庭和贫困水平.
科学领域:
- 公共卫生 公共卫生
- 数据科学数据科学数据科学
- 城市卫生 城市卫生
背景情况:
- 低出生体重 (LBW) 即使在发达国家,也构成了重大的公共卫生挑战.
- 在城市环境中对LBW率的生态/人口水平预测对于有针对性的干预至关重要.
研究的目的:
- 评估机器学习模型在美国大城市内预测LBW率方面的有效性.
- 确定与LBW率相关的关键生态/人口水平预测因素.
主要方法:
- 利用来自大城市健康库存数据平台 (2010-2022) 的公开数据.
- 分析了美国35个最大,最多城市城市的数据.
- 采用模型不可知论的方法来识别和可视化有影响力的预测因素.
主要成果:
- 机器学习模型显示出出色的预测性能 (R平方值在0.79到0.82之间).
- 最好的子集选择模型只用四个预测器实现了高精度.
- 有意义的预测因素包括克拉米迪亚感染率,种族隔离,产前护理,单亲家庭百分比和贫困.
结论:
- 机器学习算法在预测城市人口中的LBW率方面非常有效.
- 识别的预测因素为决策者和卫生当局提供了可操作的见解,以解决LBW.
相关概念视频
Regression Toward the Mean
6.5K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.5K
Steps in Outbreak Investigation
215
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
215
z Scores and Area Under the Curve
11.4K
z scores are the standardized values obtained after converting a normal distribution into a standard normal distribution. A z score is measured in units of the standard deviation. The z score tells you how many standard deviations the value x is above (to the right of) or below (to the left of) the mean, μ. Values of x that are larger than the mean have positive z scores, and values of x that are smaller than the mean have negative z scores. If x equals the mean, then x has a z score of...
11.4K
Prediction Intervals
2.4K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.4K
Mechanistic Models: Compartment Models in Individual and Population Analysis
89
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
89
Residuals and Least-Squares Property
7.9K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.9K


