评估存在缺席模型的预测性能:为什么同一个模型看起来很好或很差?
Nerea Abrego1,2, Otso Ovaskainen1,3,4
1Department of Biological and Environmental Science University of Jyväskylä Jyväskylä Finland.
Ecology and evolution
|December 19, 2023
概括
选择最好的物种分布模型需要仔细评估预测性表现指标. 不同的指标和空间尺度可以产生不同的结果,强调需要全面的方法,而不是简单的规则.
科学领域:
- 生态生态学 生态生态学
- 生物地理学 生物地理学 生物地理学
- 计算生物学 计算生物学
背景情况:
- 选择最佳物种分布模型 (SDM) 取决于预测性能.
- 确定模型的性能是否"足够好"仍然是一个重大挑战.
- 存在缺席模型对于理解物种分布至关重要.
研究的目的:
- 澄清评估SDM预测绩效的关键选择和指标.
- 评估不同性能指标与模型组件和空间尺度的关系.
- 引导研究人员解释和比较SDM预测性能.
主要方法:
- 为了评估四个预测性绩效指标,使用了一个层次化的案例研究:曲线下的面积 (AUC),Tjur的R平方,最大的Kappa (max-Kappa) 和最大的真实技能统计 (max-TSS).
- 该研究研究了随机和固定效应,空间尺度和交叉验证策略对这些指标的影响.
- 在不同的空间尺度上测量性能,以观察度量变化.
主要成果:
- 预测性绩效指标可以根据测量的空间尺度为同一模型产生不同的值.
- 图尔的R平方和最大卡帕对物种流行敏感,通常在较小的规模下降.
- AUC和最大TSS在很大程度上独立于患病率,在空间尺度上显示一致的值,但提供互补的见解.
结论:
- 模型性能解释需要谨慎,因为指标因规模和计算方法而异.
- 互补指标的组合提供了比单个措施或绝对值更全面的评估.
- 研究人员应该将模型性能与先验预期进行比较,而不是依赖简单的经验规则.
相关概念视频
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K
Expected Frequencies in Goodness-of-Fit Tests
2.5K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.5K
Mechanistic Models: Compartment Models in Individual and Population Analysis
43
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
43
Predicting Products: Substitution vs. Elimination
11.7K
When a nucleophile and an alkyl halide react, nucleophilic substitution and β-elimination reactions compete to generate products.
The following factors can influence the mechanisms competing against each other:
The following factors can influence the mechanisms competing against each other:
11.7K
Sensitivity, Specificity, and Predicted Value
385
In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
Sensitivity is the...
385
Clearance Models: Noncompartmental Models
62
Clearance is a pharmacokinetic parameter traditionally defined by compartment models, signifying the rate at which a drug is expelled from the body. However, a noncompartmental model offers an alternative method for assessing clearance, primarily employing empirical data obtained after administering a single drug dose.
The noncompartmental approach capitalizes on extensive sampling data, correlating the volume of distribution to systemic exposure and the administered dosage. This method enables...
The noncompartmental approach capitalizes on extensive sampling data, correlating the volume of distribution to systemic exposure and the administered dosage. This method enables...
62


