预测模型的低出生体重在怀孕:一个对逻辑回归和决策树方法的比较分析方法
Ravi Kumar1, Abhinav Bahuguna2, Palak Goyal1
1Department of Community Medicine, Shri Ram Murti Smarak Institute of Medical Sciences, Bareilly, India.
概括
预测低出生体重 (LBW) 对婴儿健康至关重要. 孕产妇年龄,并发症和妊娠年龄显著预测LBW,后勤回归显示的精度高于决策树.
科学领域:
- 孕产妇和儿童的健康
- 生物统计学 生物统计学
- 预测建模预测建模
背景情况:
- 出生体重对婴儿发育至关重要.
- 低出生体重 (LBW) 婴儿面临着重大的早期健康挑战.
- 识别LBW预测因素对于有针对性的干预措施至关重要.
研究的目的:
- 用基于模型的方法识别LBW的显著预测因素.
- 为了比较物流回归和决策树模型对LBW的预测性能.
主要方法:
- 130名孕妇 (2022-2023) 的医院横截面研究.
- 应用物流回归和决策树方法.
- 使用接收器运行特征 (ROC) 曲线评估模型性能.
主要成果:
- LBW的患病率为38.5%,为5%.
- 显著的LBW预测因素包括母亲的年龄,堕胎史,并发症,妊娠并发症和妊娠年龄 (P <0.05).
- 逻辑回归 (AUC=0.881) 和决策树 (AUC=0.814) 模型显示出良好的区分能力.
结论:
- 与决策树模型相比,物流回归在预测LBW方面表现出更高的准确性.
- 调查结果强调需要有针对性的母婴护理政策,以减轻LBW风险.
- 决策树虽然对模式识别有用,但由于潜在的过度拟合,需要谨慎应用.
相关概念视频
Survival Tree
389
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
389
Regression Toward the Mean
6.9K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.9K
Regression Analysis
8.1K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
8.1K
Residuals and Least-Squares Property
9.1K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
9.1K
Multiple Regression
3.8K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.8K
z Scores and Area Under the Curve
18.4K
z scores are the standardized values obtained after converting a normal distribution into a standard normal distribution. A z score is measured in units of the standard deviation. The z score tells you how many standard deviations the value x is above (to the right of) or below (to the left of) the mean, μ. Values of x that are larger than the mean have positive z scores, and values of x that are smaller than the mean have negative z scores. If x equals the mean, then x has a z score of...
18.4K

