数据平衡方法对使用基于随机森林随机参数的非侵入性参数预测代谢综合征的影响
Sahar Mohseni-Takalloo1,2,3, Hadis Mohseni4, Hassan Mozaffari-Khosravi2,3
1School of Public Health, Bam University of Medical Sciences, Bam, Iran.
BMC bioinformatics
|January 11, 2024
概括
预测代谢综合征 (MetS) 对于识别糖尿病和心血管疾病的风险至关重要. 使用非侵入性数据的随机森林模型,增强了SplitBal平衡,显著提高了男性和女性的预测灵敏度.
科学领域:
- 生物医学信息学 生物医学信息学
- 医疗保健中的机器学习
- 公共卫生查 公共卫生查
背景情况:
- 代谢综合征 (MetS) 是糖尿病和心血管疾病的关键预测因素.
- 目前的诊断往往需要生物化学测试,突出需要更简单的方法.
- 对MetS的非侵入性预测可以有效地识别有风险的人群.
研究的目的:
- 使用随机森林算法使用非侵入性特征预测代谢综合征 (MetS).
- 评估数据平衡技术 (SMOTE和SplitBal) 对MetS预测模型性能的影响.
- 为了解决MetS数据集中固有的数据不平衡.
主要方法:
- 使用随机森林学习算法进行MetS预测.
- 研究了两个数据平衡技术:合成少数人过量采样技术 (SMOTE) 和随机分割数据平衡 (SplitBal).
- 在男性和女性中使用准确度和灵敏度指标评估模型性能.
主要成果:
- 腰围成为MetS最重要的预测因素.
- 在没有数据平衡的情况下,随机森林模型的准确度达到86.9% (男性) 和79.4% (女性),但敏感度低 (37.1%和38.2%).
- 尽管精度下降,但SplitBal技术显著提高了对82.3% (男性) 和73.7% (女性) 的灵敏度.
结论:
- 随机森林与数据平衡相结合,特别是SplitBal,为MetS预测提供了一个有希望的方法.
- 这些模型可以在健康查计划中作为有价值的预后工具.
- 对于早期风险识别而言,对MetS的非侵入性预测是可行的,也是有益的.
相关概念视频
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Survival Tree
86
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
86
Strategies for Assessing and Addressing Confounding
101
Confounding is a critical issue in epidemiological studies, often leading to misleading conclusions about associations between exposures and outcomes. It occurs when the relationship between the exposure and the outcome is mixed with the effects of other factors that influence the outcome. Given that, addressing confounding is of high importance for drawing accurate inferences in research.
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
101
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
130
Statistical inference techniques, paramount in hypothesis testing, differentiate into two broad categories: parametric and nonparametric statistics.
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
130
Mechanistic Models: Compartment Models in Individual and Population Analysis
43
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
43


