使用逻辑回归和机器学习方法预测多胎妇女的早产
1Clinical Research Unit (CRU), CHEO Research Institute, University of Ottawa, Ottawa, Canada. rezaarabi11@gmail.com.
Scientific reports
|September 20, 2024
概括
在使用机器学习和逻辑回归的多胎妇女中预测早产 (PTB) 显示了类似的性能. 模型实现了高负预测值,有助于早期识别PTB风险.
科学领域:
- 围产期流行病学 围产期流行病学
- 生殖健康 生殖健康
- 临床信息学是一种临床信息学.
背景情况:
- 过早分娩 (PTB) 仍然是新生儿发病率和死亡率的主要原因.
- 准确预测多夫妇女性的PTB对于及时干预至关重要.
- 将高级机器学习 (ML) 与传统的后勤回归 (LR) 进行比较对于PTB预测至关重要.
研究的目的:
- 将ML算法的预测性能与早产 (PTB) 的后勤回归进行比较.
- 在怀孕的第一和第二个三个月期间,在多夫妇妇女中确定PTB的关键预测因子.
- 用AUC和负预测值等指标来评估这些模型的临床实用性.
主要方法:
- 一项基于人口的队列研究,使用了来自安大略省更好的结果注册和网络 (BORN) 的数据.
- 逐步后勤回归和Boruta算法用于第一和第二季度的变量选择.
- 模型的训练和验证包括LR,随机森林,决策树和人工神经网络;使用过量采样进行平衡;通过十倍交叉验证进行优化.
主要成果:
- 该队列包括145,846名分娩,其中5.57%为早产.
- 第一个三个月的模型确定了以前的PTB,糖尿病和异常怀孕相关的血蛋白-A作为关键预测因素;ANN达到68.8%的AUC.
- 第二个三个月的模型,包括婴儿性别和并发症等额外的变量,改善了LR的AUC到80.5%;所有模型都实现了~97%的负预测值.
结论:
- 机器学习和逻辑回归都在预测早产时表现相似.
- 包含妊娠并发症的第二个三个月的模型显著提高了PTB预测的准确性.
- 这些模型的高负预测值 (~97%) 表明它们在排除PTB风险方面的潜在实用性.
更多相关视频
相关概念视频
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Probability Laws
40.7K
Overview
40.7K
Residuals and Least-Squares Property
7.3K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.3K


