开发一种诊断预测模型,用于确定马拉维儿童的发展迟缓:对变量选择方法的比较分析
Jonathan Mkungudza1, Halima S Twabi2, Samuel O M Manda3
1Department of Mathematical Sciences, University of Malawi, Zomba, Malawi.
BMC medical research methodology
|August 8, 2024
概括
用各种可变选择方法比较了儿童发育迟缓预测模型. 判断方法确定了关键的风险因素,产生了一个有用的模型,用于早期识别有风险的儿童进行营养干预.
科学领域:
- 营养科学 营养科学
- 公共卫生 公共卫生
- 生物统计学 生物统计学
背景情况:
- 儿童发育迟缓是营养不良的关键指标,也是全球卫生优先事项.
- 儿童发育迟缓的预测模型需要仔细选择预测变量以获得最佳的性能.
- 了解风险因素对于开发有效的干预措施至关重要.
研究的目的:
- 为了比较儿童发育迟缓的诊断预测模型的性能.
- 评估不同的变量选择方法来识别关键的发育迟缓预测因素.
- 开发一个风险评分来识别高风险的儿童.
主要方法:
- 文献审查,以确定撒哈拉以南非洲的衰老决定因素.
- 使用马拉维人口健康调查 (MDHS 2015-16) 数据进行多变量逻辑回归.
- 应用七种可变选择算法 (向后,向前,逐步,随机森林,LASSO,判断) 来识别预测因素.
- 用AUROC,灵敏度和特异性计算儿童衰退风险得分和评估.
主要成果:
- 确定了68个潜在的预测变量,其中27个可在数据集中使用.
- 通常选择的风险因素包括家庭财富,孩子的年龄,家庭规模,出生类型和出生体重.
- 判断方法产生了最好的风险预测模型,在测试数据中,接收器运营商曲线下的面积 (AUROC) 为64%.
- 城市 (67%) 的AUROC略高于农村 (63%) 的儿童.
结论:
- 开发的儿童缩诊断预测模型可以作为初始查工具.
- 早期识别有风险的儿童有助于及时进行营养干预.
- 变量选择方法显著影响滞后预测模型的性能.
相关概念视频
Survival Tree
74
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
74
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
47
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
47
Mechanistic Models: Compartment Models in Individual and Population Analysis
35
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
35
Residuals and Least-Squares Property
7.3K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.3K
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K


