Related Experiment Video
Updated: May 20, 2025

Strand-Specific Analysis of Proteins at Replicating DNA Strands by Enrichment and Sequencing of Protein-Associated Nascent DNA Method
Published on: May 2, 2025
Correcting for Bias in Estimates of θ w and Tajima's D From Missing Data in Next-Generation Sequencing
Nick Bailey1,2, Laurie Stevison2, Kieran Samuk3
1Laboratory of Biometry and Evolutionary Biology, University of Lyon 1, CNRS UMR 5558, Lyon, France.
Missing genomic data can skew population genetic analyses. This study shows how missing data biases estimates of genetic diversity and evolution, but provides corrected methods in pixy software to improve accuracy.
Area of Science:
- Population genetics
- Genomics
- Evolutionary biology
Background:
- Population genetic analyses rely on the site frequency spectrum to understand evolutionary processes.
- Watterson's estimator (θ) and Tajima's D are key statistics for genetic diversity and detecting non-neutral evolution.
- Missing data in genomic datasets, especially in Variant Call Format (VCF) files, can introduce biases into these estimates.
Purpose of the Study:
- To assess the impact of missing genomic data on population genetic summary statistics.
- To evaluate biases in Watterson's estimator and Tajima's D across different software.
- To develop and implement methods for correcting these biases.
Main Methods:
- Simulated neutral genomic data with controlled levels of missing genotypes and sites.
- Analyzed simulated data using multiple population genetics software packages (VCFtools, PopGenome, pegas, scikit-allel).
- Developed and integrated bias correction functions into the pixy software.
Main Results:
- Consistent underestimation of Watterson's estimator (θ) was observed across software due to missing data.
- Biases in Tajima's D estimates were detected, with directions varying by software.
- Implemented correction methods in pixy significantly reduced the observed biases.
Conclusions:
- Missing data in population genomics can lead to erroneous evolutionary inferences.
- Accurate handling and correction of missing data are crucial for reliable population genetic studies.
- The updated pixy software provides a valuable tool for mitigating biases in population genetic analyses.
Related Concept Videos
Bias
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
Bias in Epidemiological Studies
What are Estimates?
The estimate for the mean of a sample is denoted by ͞x, whereas the mean of the population is designated as μ. Further, parameters such...
Systematic Error: Methodological and Sampling Errors
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Contaminants and Errors
Another key consideration is determining the appropriate number of samples required to...

