医学中随机森林变量重要性指标的使用和滥用:通过事件中风预测的演示
Meredith L Wallace1,2, Lucas Mentch3, Bradley J Wheeler4
1Department of Psychiatry, University of Pittsburgh, 3811 O'Hara Street, Pittsburgh, PA, 15231, USA. lotzmj@upmc.edu.
BMC medical research methodology
|June 19, 2023
概括
研究人员经常在随机森林模型中使用有偏见的"自制"变量重要性指标 (VIMPs). 现代的"淘汰VIMP"为分析复杂的医疗数据提供了更易于解释和更准确的替代方案.
科学领域:
- 医疗信息学医学信息学
- 统计建模 统计建模
- 机器学习是机器学习.
背景情况:
- 随机森林对于分析复杂的医疗数据来说非常强大.
- 目前的"包外" (OOB) 变量重要性指标 (VIMPs) 有统计上的局限性,包括偏差和糟糕的解释性.
- 一个现代的替代方案,即"淘汰VIMPs",为理解模型预测提供了优势.
研究的目的:
- 评估当前OOB VIMP在医学研究中的局限性.
- 引入和推使用可解释的"敲门VIMP"的有组织策略.
- 在预测事故中风时,比较OOB和伪造VIMP.
主要方法:
- 对最近50个随机森林手稿的文献综述,以评估VIMP使用情况.
- 开发一种随机森林模型,利用睡眠心脏健康研究数据预测5年内发生的中风.
- 使用OOB VIMP与仿制VIMP获得的结果的比较.
主要成果:
- 几乎所有审查的论文都在VIMP应用中表现出重大局限性.
- OOB VIMPs确定了相关的肺功能变量作为关键中风预测因素.
- 伪造的VIMP揭示了药物,医疗风险因素,年龄,血压和吸烟作为重要的预测因素,导致了不同的结论.
结论:
- 过度依赖OOB VIMP可能会产生误导性的研究结果.
- 实现可解释和无偏见的"淘汰VIMP"对于推进医学机器学习应用至关重要.
- 假冒VIMP的广泛采用将引导研究人员进行更有意义的发现.
相关概念视频
Biostatistics: Overview
287
Biostatistics plays a crucial role in understanding and analyzing data in healthcare and biology. Biostatisticians conduct experiments, gather evidence, and draw meaningful conclusions using statistical methods and techniques. Different variables form the foundation of biostatistical analysis, allowing researchers to understand and interpret data effectively. These variables are classified into different types, each serving a specific purpose in statistical analysis.
Discrete variables are...
Discrete variables are...
287
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Confounding in Epidemiological Studies
198
Confounding in statistical epidemiology represents a pivotal challenge, referring to the distortion in the perceived relationship between an exposure and an outcome due to the presence of a third variable, known as a confounder. This variable is associated with both the exposure and the outcome but is not a direct link in their causal chain. Its presence can lead to erroneous interpretations of the exposure's effect, either exaggerating or underestimating the true association. This...
198
Statistical Methods for Analyzing Epidemiological Data
427
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
427
Bias in Epidemiological Studies
380
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:
380
Sensitivity, Specificity, and Predicted Value
513
In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
Sensitivity is the...
513


