Related Experiment Video
Updated: Jun 17, 2026

14:06
Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
Published on: June 23, 2012
A SNP discovery method to assess variant allele probability from next-generation resequencing data
Yufeng Shen1, Zhengzheng Wan, Cristian Coarfa
1The Human Genome Sequencing Center, Baylor College of Medicine, Houston, Texas 77030, USA. yshen@c2b2.columbia.edu
Genome Research
|December 19, 2009
Summary
Atlas-SNP2 is a new tool that identifies genetic variants from next-generation sequencing (NGS) data. It distinguishes true single nucleotide polymorphisms (SNPs) from sequencing errors using a logistic regression model and Bayesian analysis.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Accurate genetic variant identification is crucial for large-scale genomic projects like the 1000 Genomes Project.
- Distinguishing true single nucleotide polymorphisms (SNPs) from sequencing errors is a major challenge in next-generation sequencing (NGS) data analysis.
- Understanding base call error probabilities is essential for reliable SNP discovery.
Purpose of the Study:
- To develop a computational tool, Atlas-SNP2, for accurate SNP detection from NGS data.
- To account for systematic sequencing errors influenced by context-related variables.
- To estimate the posterior error probability for substitutions to differentiate true SNPs from errors.
Main Methods:
- Developed Atlas-SNP2, a computational tool utilizing a logistic regression model trained on datasets.
- Incorporated context-related variables to detect and account for systematic sequencing errors.
- Employed a Bayesian formula to estimate posterior error probabilities, integrating prior error rates and SNP rates.
Main Results:
- Atlas-SNP2 effectively distinguishes true SNPs from sequencing errors.
- Achieved a false-positive rate below 10%.
- Demonstrated a false-negative rate of approximately 5% or lower.
Conclusions:
- Atlas-SNP2 provides a robust method for accurate SNP identification in NGS data.
- The tool's ability to account for systematic errors improves variant calling reliability.
- The low false-positive and false-negative rates make Atlas-SNP2 valuable for genomic research.
Related Concept Videos
Comparing Copy Number Variations and SNPs
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Genome-wide Association Studies-GWAS
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
Single Nucleotide Polymorphisms-SNPs
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
Next-generation Sequencing
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features.

