罕见变异分析中的获胜者诅咒:影响大小估计偏差取决于影响方向和使用的关联方法
David Soave1,2, Melisa Hayalioglu1, Lei Sun3,4
1Department of Mathematics, Wilfrid Laurier University, Waterloo, ON, Canada.
Frontiers in genetics
|August 27, 2025
概括
由于偏见,对复杂特征的罕见变异 (RV) 影响的估计具有挑战性. 这项研究表明,获胜者的诅咒和变异性如何在平均遗传效应 (AGE) 估计中产生竞争的上下偏差.
科学领域:
- 遗传学
- 统计遗传学
- 人类的复杂特征
背景情况:
- 很大一部分复杂的人类特征的遗传性并不能通过常见的变异来解释.
- 估计罕见变异对复杂特征病因的贡献至关重要.
- 基于基因的RV测试方法将多个变体的信息结合起来,以检测关联,但个体效应估计受到样本大小的限制.
研究的目的:
- 调查平均遗传效应 (AGE) 和罕见变异 (RV) 的个体变异效应的竞争上下偏差.
- 用偏差校正技术说明这些偏差对聚合估计的精度的影响.
- 检查个人因果变异效应如何在RV被组合时导致偏差.
主要方法:
- 进行了一项模拟研究,以模拟获胜者诅咒和异质性对变异效应大小估计的影响.
- 该研究评估了各种偏差校正技术的性能,包括重新抽样和基于概率的方法.
- 在模拟复制品中分析了因果变异的个体效应估计.
主要成果:
- 平均遗传效应 (AGE) 和个体变异效应都受到竞争的上升 (获胜者诅咒) 和下降 (异质性) 偏差的影响.
- 这些相互竞争的偏差可能会使偏差校正技术复杂化,从而影响聚合估计的准确性.
- 当罕见变异 (RV) 聚合时,个别变异效应估计有助于观察到的偏差.
结论:
- 了解和纠正竞争偏差对于精确估计复杂特征的罕见变异 (RV) 效应至关重要.
- 单个变异效应的异质性和胜利者的诅咒在统计遗传学中提出了重大挑战.
- 模拟研究对于剖析罕见变异关联研究中偏差的复杂相互作用是有价值的.
相关概念视频
Bias
4.9K
Bias refers to any tendency that prevents a question from being considered unprejudiced. In research, bias occurs when one outcome or answer is selected or encouraged over others in sampling or testing. Bias can occur during any research phase, including study design, data collection, analysis, and publication.
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
4.9K
Cause and Effect
11.3K
While variables are sometimes correlated because one does cause the other, it could also be that some other factor, a confounding variable, is actually causing the systematic movement in our variables of interest. For instance, as sales in ice cream increase, so does the overall rate of crime. Is it possible that indulging in your favorite flavor of ice cream could send you on a crime spree? Or, after committing crime do you think you might decide to treat yourself to a cone?
11.3K
Regression Toward the Mean
6.5K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.5K
Correlation and Causation
39.5K
Statistical tests can calculate whether there is a relationship, or correlation, between independent and dependent variables. An indirect relationship of the variables signifies a correlation, while a direct relationship shows causation. If it is determined that no connection exists between the variables, then the correlation is a coincidence.
Correlation versus Causation
If the dependent variable increases or decreases when the independent variable increases, there is a positive or negative...
Correlation versus Causation
If the dependent variable increases or decreases when the independent variable increases, there is a positive or negative...
39.5K
Wilcoxon Rank-Sum Test
342
The Wilcoxon rank-sum test, also known as the Mann-Whitney U test, is a nonparametric test used to determine if there is a significant difference between the distributions of two independent samples. This test is designed specifically for two independent populations and has the following key requirements:
342
Friedman Two-way Analysis of Variance by Ranks
296
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
296


