相关实验视频
Updated: Jul 18, 2025

05:37
An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
2.1K
机器学习模型的配对评估描述了混因素和异常值的影响
Maulik K Nariya1,2, Caitlin E Mills1, Peter K Sorger1,2
1Laboratory of Systems Pharmacology, Harvard Program in Therapeutic Science, Harvard Medical School, Boston, MA 02115, USA.
Patterns (New York, N.Y.)
|August 21, 2023
概括
准确的机器学习模型评估至关重要. 配对评估提供了一种可靠的方法,用于评估小型生物和临床研究中的预测器性能,提高准确性和识别异常值.
科学领域:
- 机器学习 机器学习
- 生物统计学 生物统计学
- 计算生物学 计算生物学
- 临床研究 临床研究
背景情况:
- 真正的机器学习模型准确度是人口层面的统计数据,不能直接观察到.
- 估计预测器性能依赖于测试数据集,其代表性影响准确性.
- 在生物学和医学中,小样本研究在可靠的模型评估方面面临挑战.
研究的目的:
- 引入配对评估,一种简单而强大的方法来评估机器学习模型的性能.
- 评估测试数据选择对小样本研究中的性能估计的影响.
- 在生物和临床应用中证明配对评估的实用性.
主要方法:
- 机器学习模型评估对对评估技术的描述.
- 配对评估的应用,以预测乳腺癌细胞系中的药物反应.
- 配对评估的应用,以预测阿尔茨海默病患者的严重程度.
主要成果:
- 测试数据的选择可能导致性能估计变化高达20%.
- 配对评估有效地识别了绩效估计中的异常值.
- 该方法在存在混因素时提高了准确性,并允许对模型比较进行统计显著性赋值.
结论:
- 配对评估为在小样本环境中对机器学习模型性能评估提供了可靠的方法.
- 这种方法提高了在关键的生物和临床研究中模型评估的可靠性.
- 配对评估有助于更准确地比较和解释机器学习模型.
相关概念视频
Confounding in Epidemiological Studies
188
Confounding in statistical epidemiology represents a pivotal challenge, referring to the distortion in the perceived relationship between an exposure and an outcome due to the presence of a third variable, known as a confounder. This variable is associated with both the exposure and the outcome but is not a direct link in their causal chain. Its presence can lead to erroneous interpretations of the exposure's effect, either exaggerating or underestimating the true association. This...
188
Outliers and Influential Points
4.1K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
4.1K
Strategies for Assessing and Addressing Confounding
119
Confounding is a critical issue in epidemiological studies, often leading to misleading conclusions about associations between exposures and outcomes. It occurs when the relationship between the exposure and the outcome is mixed with the effects of other factors that influence the outcome. Given that, addressing confounding is of high importance for drawing accurate inferences in research.
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
119
Quantifying and Rejecting Outliers: The Grubbs Test
1.7K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.7K
What Are Outliers?
3.9K
Outliers are observed data points that are far from the least squares line. They have unusual values and need to be examined carefully. Though an outlier may result from erroneous data, at other times, it may hold valuable information about the population under study and should be included in the data. Hence, it is crucial to examine what causes a data point to be an outlier.
The z score is used to find outliers or unusual values. It should be noted that any values beyond -2 and +2 are...
The z score is used to find outliers or unusual values. It should be noted that any values beyond -2 and +2 are...
3.9K
Detection of Gross Error: The Q Test
6.2K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.2K

