Related Experiment Video
Updated: May 8, 2026

10:36
Rare Event Detection Using Error-corrected DNA and RNA Sequencing
Published on: August 3, 2018
Estimating inbreeding coefficients from NGS data: Impact on genotype calling and allele frequency estimation
Filipe G Vieira1, Matteo Fumagalli, Anders Albrechtsen
1Department of Integrative Biology, University of California, Berkeley, Berkeley, California 94720, USA;
Genome Research
|August 17, 2013
Summary
This study introduces two methods for estimating inbreeding coefficients from next-generation sequencing (NGS) data, improving genotype accuracy in highly inbred species with low-coverage data.
Area of Science:
- Population genetics
- Genomics
- Bioinformatics
Background:
- Next-generation sequencing (NGS) analyses often assume Hardy-Weinberg equilibrium (HWE), which is frequently violated in domesticated or selfing species.
- Deviations from HWE are common in many organisms, particularly those with asexual life cycles or undergoing domestication.
- Accurate genotype calling and site frequency spectrum (SFS) estimation are challenging for low-coverage data in non-HWE populations.
Purpose of the Study:
- To develop and validate novel methods for estimating individual inbreeding coefficients (F) from NGS data.
- To assess the impact of incorporating inbreeding information on genotype calling and SFS estimation.
- To improve the accuracy of genomic analyses for species deviating from HWE, especially with low-coverage data.
Main Methods:
- Developed two novel methods utilizing an expectation-maximization (EM) algorithm to estimate inbreeding coefficients (F).
- Applied these methods to both simulated and real next-generation sequencing (NGS) datasets.
- Evaluated the performance of genotype calling and site frequency spectrum (SFS) estimation with and without accounting for inbreeding.
Main Results:
- The proposed methods accurately estimate inbreeding coefficients (F) from NGS data, even with low coverage.
- Accounting for inbreeding significantly increases the accuracy of genotype calling and SFS estimation in highly inbred samples.
- Demonstrated marked improvements in accuracy for low-coverage, highly inbred datasets using the new methods.
Conclusions:
- The developed expectation-maximization (EM) based methods provide accurate inbreeding coefficient (F) estimates from NGS data.
- Incorporating inbreeding information enhances the reliability of genotype calling and SFS estimation in non-HWE populations.
- These methods are effective for improving genomic analyses in species with significant inbreeding, especially when using low-coverage sequencing data.
Related Concept Videos
Mutation, Gene Flow, and Genetic Drift
In a population that is not at Hardy-Weinberg equilibrium, the frequency of alleles changes over time. Therefore, any deviations from the five conditions of Hardy-Weinberg equilibrium can alter the genetic variation of a given population. Conditions that change the genetic variability of a population include mutations, natural selection, non-random mating, gene flow, and genetic drift (small population size).
Comparing Copy Number Variations and SNPs
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
