在下一代测序中的缺失数据中纠正估计的偏差和Tajima的偏差
Nick Bailey1,2, Laurie Stevison2, Kieran Samuk3
1Laboratory of Biometry and Evolutionary Biology, University of Lyon 1, CNRS UMR 5558, Lyon, France.
Molecular ecology resources
|March 24, 2025
概括
缺少基因组数据可能会扭曲人口遗传分析. 这项研究表明,缺少的数据如何对遗传多样性和进化的估计产生偏见,但在pixy软件中提供了更正的方法来提高准确性.
科学领域:
- 人口遗传学 人口遗传学
- 基因组学就是基因组学.
- 进化生物学是进化的生物学.
背景情况:
- 种群遗传分析依赖于现场频谱来理解进化过程.
- 沃特森的估计量 (θ) 和塔吉马的D是遗传多样性和检测非中性进化的关键统计数据.
- 基因组数据集中缺少的数据,特别是变异调用格式 (VCF) 文件中缺少的数据,可以在这些估计中引入偏差.
研究的目的:
- 评估缺少的基因组数据对人口遗传总结统计数据的影响.
- 为了在不同的软件中评估沃特森的估计器和塔吉马的D中的偏差.
- 开发和实施纠正这些偏见的方法.
主要方法:
- 模拟的中性基因组数据与失踪基因型和遗址的控制水平.
- 使用多个种群遗传学软件包 (VCFtools,PopGenome,pegas,scikit-allel) 分析模拟数据.
- 在pixy软件中开发并集成了偏差校正功能.
主要成果:
- 由于缺少数据,在软件中观察到对沃特森估计器 (θ) 的持续低估.
- 在Tajima的D估计中检测到偏差,方向因软件而异.
- 在pixy中实施的校正方法显著减少了观察到的偏差.
结论:
- 人口基因组学中缺少的数据可能导致错误的进化推断.
- 准确处理和纠正缺失的数据对于可靠的人口遗传研究至关重要.
- 更新后的pixy软件为减轻人口遗传分析中的偏见提供了有价值的工具.
相关概念视频
Bias
3.7K
Bias refers to any tendency that prevents a question from being considered unprejudiced. In research, bias occurs when one outcome or answer is selected or encouraged over others in sampling or testing. Bias can occur during any research phase, including study design, data collection, analysis, and publication.
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
3.7K
Bias in Epidemiological Studies
112
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:
112
What are Estimates?
4.9K
It isn't easy to measure a parameter such as the mean height or the mean weight of a population. So, we draw samples from the population and calculate the mean height or mean weight of the individuals in the sample. This sample data acts as a representative measure of the population parameter. These sample statistics are known as estimates.
The estimate for the mean of a sample is denoted by ͞x, whereas the mean of the population is designated as μ. Further, parameters such...
The estimate for the mean of a sample is denoted by ͞x, whereas the mean of the population is designated as μ. Further, parameters such...
4.9K
Systematic Error: Methodological and Sampling Errors
1.4K
In the case of systematic errors, the sources can be identified, and the errors can be subsequently minimized by addressing these sources. According to the source, systematic errors can be divided into sampling, instrumental, methodological, and personal errors.
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
1.4K
Strategies for Assessing and Addressing Confounding
77
Confounding is a critical issue in epidemiological studies, often leading to misleading conclusions about associations between exposures and outcomes. It occurs when the relationship between the exposure and the outcome is mixed with the effects of other factors that influence the outcome. Given that, addressing confounding is of high importance for drawing accurate inferences in research.
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
77
Contaminants and Errors
81
Effective sample preparation is crucial for accurate and reliable laboratory analysis. During this process, two significant sources of error can arise: concentration bias from improper sample splitting and contamination caused by methods used to reduce particle size, such as grinding or homogenization. Identifying and minimizing these potential errors is crucial to ensuring the validity of the analysis.
Another key consideration is determining the appropriate number of samples required to...
Another key consideration is determining the appropriate number of samples required to...
81


