在积极未标记的环境中检测预测模型的偏见验证:疾病基因优先级案例研究
Ivan Molotkov1,2,3, Mykyta Artomov1,2
1The Steve and Cindy Rasmussen Institute for Genomic Medicine, Nationwide Children's Hospital, Columbus, OH, United States.
Bioinformatics advances
|September 25, 2023
概括
验证积极未标记模型的完全随机 (SCAR) 假设经常被违反,导致膨胀的性能估计. 我们的算法检测到这种验证偏差,对于复杂的遗传特征中准确的基因优先级来说至关重要.
科学领域:
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
- 遗传学 遗传学 是一个
背景情况:
- 在医学,遗传学和生物研究中,积极的未标记数据是常见的,需要准确的预测模型.
- 模型性能通常使用被认为完全随机从已知的积极示例中选择的数据集 (SCAR) 进行验证.
- 虽然SCAR假设很方便,但在实践中往往缺乏严格的理由.
研究的目的:
- 开发和验证一种算法,用于检测在积极未标记的学习中违反SCAR假设的情况.
- 调查复杂遗传特征基因优先排序中的验证偏差的流行程度和影响.
- 为生物信息学预测模型提供一个更可靠的性能估计工具.
主要方法:
- 开发了一种算法来测试SCAR假设违反在最小假设下 (积极示例的下限).
- 将算法应用于复杂遗传特征的基因优先级数据集.
- 量化了验证偏差的程度及其对绩效指标的影响.
主要成果:
- 在基因优先级数据中,SCAR假设经常被违反,导致显著的性能高估 (验证偏差).
- 验证偏差可以膨胀性能估计,解释报告的模型有效性和实际实用性之间的差异.
- 开发的算法有效地识别和量化这种偏差.
结论:
- 该SCAR假设不是普遍适用的,其违反引入了积极的未标记模型评估的实质性偏见.
- 验证偏差是基因优先级的普遍问题,可能会误导模型性能.
- 准确的模型评估需要仔细考虑和测试SCAR假设,特别是在复杂的生物数据分析中.
相关概念视频
Sensitivity, Specificity, and Predicted Value
481
In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
Sensitivity is the...
481
Bias in Epidemiological Studies
339
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:
339
Genetic Screens
5.0K
Genetic screens are tools used to identify genes and mutations responsible for phenotypes of interest. Genetic screens help identify individuals or a group of people at risk of developing genetic diseases and help them with early intervention, targeted therapy, and reproductive options.
Forward genetic screens
Forward or “classical” genetic screens involve creating random mutations in an organism’s DNA using radiation, mutagens, or insertion of additional bases, which...
Forward genetic screens
Forward or “classical” genetic screens involve creating random mutations in an organism’s DNA using radiation, mutagens, or insertion of additional bases, which...
5.0K
Genome-wide Association Studies-GWAS
13.6K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
13.6K
Data Validation
5.1K
Data validation is an essential part of a comprehensive assessment. Validation is confirming or verifying and opening the door to gathering more assessment data as it clarifies vague or unclear data. The process of checking and verifying the collected information is called data validation. The primary purpose of data validation is to ensure data is as free from error, bias, and misinterpretation as possible.
Nursing assessment guides are generally based on holistic models rather than medical...
Nursing assessment guides are generally based on holistic models rather than medical...
5.1K


