Related Experiment Video
Updated: May 22, 2026

08:38
Targeted DNA Methylation Analysis by Next-generation Sequencing
Published on: February 24, 2015
Association testing for next-generation sequencing data using score statistics
Line Skotte1, Thorfinn Sand Korneliussen, Anders Albrechtsen
1Department of Biology, University of Copenhagen, Copenhagen, Denmark. line@binf.ku.dk
Genetic Epidemiology
|May 10, 2012
Summary
This study introduces a computationally feasible score statistic for genetic association testing in large sequencing studies. It improves power by accounting for genotype uncertainty, outperforming methods using only called genotypes.
Area of Science:
- Genomics
- Statistical Genetics
- Bioinformatics
Background:
- Large-scale sequencing studies are crucial for identifying genetic variants linked to diseases.
- Genotype calling uncertainty in sequencing data can reduce statistical power and introduce false signals.
- Existing methods to address genotype uncertainty are often computationally intensive for large datasets.
Purpose of the Study:
- To develop a computationally feasible statistical method for genetic association testing in next-generation sequencing data.
- To improve the power of association studies by explicitly accounting for genotype classification uncertainty.
- To provide a flexible framework applicable to both quantitative and discrete phenotypes, including case-control studies.
Main Methods:
- Utilizes a score statistic for the joint likelihood of observed phenotypes and sequencing data.
- Incorporates genotype uncertainty using posterior probabilities of genotypes derived from sequencing data.
- Employs a generalized linear model framework to accommodate quantitative and discrete phenotypes, allowing for covariates.
Main Results:
- The proposed score statistic method demonstrates higher statistical power compared to methods relying solely on called genotypes.
- The approach effectively accounts for genotype classification uncertainty, mitigating spurious association signals.
- The method is computationally feasible for large-scale sequencing studies.
Conclusions:
- The score statistic approach offers a powerful and computationally efficient solution for association testing with next-generation sequencing data.
- This method enhances the reliability of genetic association studies by properly handling genotype uncertainty.
- The generalized linear model framework ensures broad applicability across various study designs and phenotype types.
Related Concept Videos
Next-generation Sequencing
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features.
Modern Molecular Taxonomy
Advancements in molecular biology have revolutionized the identification and characterization of bacteria, with multiple methods leveraging DNA sequencing for enhanced precision. As sequencing technologies improve and costs decline, these approaches are increasingly used in clinical, environmental, and evolutionary studies.Multilocus Sequence Typing (MLST) examines several housekeeping genes, essential chromosomal genes encoding cellular functions, to distinguish strains. Approximately...
Genome Annotation and Assembly
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
RNA-seq
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Comparing Copy Number Variations and SNPs
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Genome-wide Association Studies-GWAS
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
