缺失项目数据对纠正风险评估工具相对预测准确性的影响
Bronwen Perley-Robertson1, Kelly M Babchishin1, L Maaike Helmus2
1Carleton University, Ottawa, Ontario, Canada.
Assessment
|February 7, 2024
概括
处理缺少的风险评估数据:比例是有效的,可与多重归算进行预测准确性相比较. 总结可用的项目低估了绝对风险,使得分数计算成为样本中缺少数据的合理方法.
科学领域:
- 犯罪学 犯罪学
- 心理测量 心理测量 心理测量
- 统计 统计 统计 统计
背景情况:
- 缺失的数据在风险评估工具中很常见.
- 缺失数据对预测准确性的影响还没有得到充分研究.
- 经常使用比较简单的方法,如总和或分数,但多重归因在理论上是优越的.
研究的目的:
- 为了比较多重归算,总和和分数的有效性,用于处理风险评估中缺少的数据.
- 研究不同百分比的缺失数据对预测准确性的影响.
主要方法:
- 使用了STABLE-2007 (N = 4,286) 和SARA-V2 (N = 455) 数据集,来自加拿大男性的社区监督.
- 引入了六个条件的缺失数据,范围从1%到50%的删除.
- 对比了三个归算技术的相对预测准确性:加法,分数和多重归算.
主要成果:
- 相对预测准确度没有受到缺失数据数量的显著影响.
- 在预测准确性方面,分量和多重归算的性能相对相似.
- 总结可用项目导致绝对风险的低估.
结论:
- 比例化是一种经验证明的有效方法,用于处理风险评估样本中缺少的数据.
- 简单的技术,比如分数计算,对于保持预测准确性,可以与多重归算一样有效.
- 风险评估从业人员在面对缺少数据时可以自信地使用分量.
更多相关视频
06:55Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
14.5K
12:18A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
7.5K
相关概念视频
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Detection of Gross Error: The Q Test
6.1K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.1K
Systematic Error: Methodological and Sampling Errors
1.5K
In the case of systematic errors, the sources can be identified, and the errors can be subsequently minimized by addressing these sources. According to the source, systematic errors can be divided into sampling, instrumental, methodological, and personal errors.
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
1.5K
Censoring Survival Data
95
Survival analysis is a statistical method used to analyze time-to-event data, often employed in fields such as medicine, engineering, and social sciences. One of the key challenges in survival analysis is dealing with incomplete data, a phenomenon known as "censoring." Censoring occurs when the event of interest (such as death, relapse, or system failure) has not occurred for some individuals by the end of the study period or is otherwise unobservable, and it might have many different...
95
Mechanistic Models: Compartment Models in Individual and Population Analysis
43
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
43
Reliability and Validity
12.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.7K
