Related Experiment Video
Updated: Apr 12, 2026

10:36
Rare Event Detection Using Error-corrected DNA and RNA Sequencing
Published on: August 3, 2018
12.7K
Evaluating information content of SNPs for sample-tagging in re-sequencing projects
Hao Hu1, Xiang Liu2, Wenfei Jin3
1Department of human molecular genetics, Max-Planck Institute for Molecular Genetics, Berlin, 14195, Germany.
Scientific Reports
|May 16, 2015
Summary
Accurate sample identification in re-sequencing is crucial. This study develops an optimized SNP panel for reliable sample tagging, requiring only 30 SNPs to distinguish 100,000 individuals, enhancing study integrity.
Area of Science:
- Genomics
- Bioinformatics
Background:
- Accidental sample mix-ups are a significant challenge in re-sequencing studies.
- Reliable sample identification is essential for data integrity and reproducibility.
Purpose of the Study:
- To develop a model for measuring Single Nucleotide Polymorphism (SNP) information content.
- To optimize SNP panels for maximal discrimination and accurate sample tagging.
Main Methods:
- Developed a computational model to quantify the information content of SNPs.
- Optimized SNP panels for efficient individual discrimination.
- Simulated large populations to test the robustness of the SNP tagging strategy.
Main Results:
- As few as 60 optimized SNPs can differentiate individuals in a global population.
- Approximately 30 optimized SNPs are sufficient to label up to 100,000 individuals.
- Simulations showed high average Hamming distances (>18) and low duality frequencies (<1 in 10,000) with 30 SNPs.
Conclusions:
- The optimized SNP panel strategy provides robust and cost-effective sample discrimination for re-sequencing.
- The developed SNP selection program allows customization for specific research needs, including Whole Exome Sequencing.
- This approach significantly improves the reliability of large-scale re-sequencing projects.
Related Concept Videos
RNA-seq
12.6K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
12.6K
Single Nucleotide Polymorphisms-SNPs
20.0K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
20.0K
Comparing Copy Number Variations and SNPs
19.4K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
19.4K
Next-generation Sequencing
101.9K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
101.9K
Sanger Sequencing
780.1K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
780.1K
Modern Molecular Taxonomy
844
Advancements in molecular biology have revolutionized the identification and characterization of bacteria, with multiple methods leveraging DNA sequencing for enhanced precision. As sequencing technologies improve and costs decline, these approaches are increasingly used in clinical, environmental, and evolutionary studies.Multilocus Sequence Typing (MLST) examines several housekeeping genes, essential chromosomal genes encoding cellular functions, to distinguish strains. Approximately...
844

