Related Experiment Video
Updated: Feb 27, 2026

05:51
A Strategy to Identify de Novo Mutations in Common Disorders such as Autism and Schizophrenia
Published on: June 15, 2011
26.5K
Optimal sequencing strategies for identifying disease-associated singletons
Sara Rashkin1,2, Goo Jun1,3, Sai Chen1,4
1Center for Statistical Genetics, Department of Biostatistics, University of Michigan, Ann Arbor, Michigan, United States of America.
Plos Genetics
|June 23, 2017
Summary
Identifying rare variants requires optimal sequencing depth. A coverage of 15-20x balances singleton variant detection power and sample size for genetic association studies of complex diseases.
Area of Science:
- Genetics
- Genomic sequencing
- Statistical genetics
Background:
- Rare variants are increasingly important in genetic association studies.
- Sequencing strategies must balance variant detection with sample size and cost-effectiveness.
- Singleton variants, found in only one individual, are challenging to detect at low sequencing depths.
Purpose of the Study:
- To investigate the impact of sequencing depth on the detection and association of rare variants, specifically singletons.
- To determine the optimal sequencing coverage for maximizing the power of rare variant association studies.
Main Methods:
- Sensitivity analysis of singleton variant detection using simulated and down-sampled deep sequencing data.
- Power calculations for case-control association studies comparing singleton variant burden.
- Evaluation across various genetic and epidemiological parameters.
Main Results:
- Singleton detection sensitivity increases with sequencing coverage, plateauing around 25x.
- Maximum association study power for singletons is achieved at 15-20x coverage when total sequencing capacity is fixed.
- Optimal coverage is independent of relative risk, disease prevalence, singleton burden, and case-control ratio.
Conclusions:
- A sequencing depth of 15-20x offers a practical compromise for rare variant studies.
- This coverage level optimizes the balance between detecting rare variants and maintaining adequate sample size.
- Informs cost-effective study design for identifying genetic associations with complex diseases.
More Related Videos
Related Concept Videos
Genome-wide Association Studies-GWAS
16.0K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
16.0K
Single Nucleotide Polymorphisms-SNPs
18.8K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
18.8K
RNA-seq
12.2K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
12.2K
Sanger Sequencing
776.0K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
776.0K
Next-generation Sequencing
99.5K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
99.5K

