Related Experiment Video
Updated: Oct 31, 2025

04:52
Following the Dynamics of Structural Variants in Experimentally Evolved Populations
Published on: February 3, 2023
1.1K
NPSV: A simulation-driven approach to genotyping structural variants in whole-genome sequencing data
Michael D Linderman1, Crystal Paudyal1, Musab Shakeel1
1Department of Computer Science, Middlebury College, 14 Old Chapel Road, Middlebury, VT 05753, USA.
Gigascience
|July 1, 2021
Summary
NPSV, a new machine learning tool, accurately genotypes structural variants (SVs) in genome sequencing data by simulating sequencing biases. This improves disease gene discovery by enhancing variant accuracy.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Structural variants (SVs) are crucial in disease etiology.
- Accurate genotyping of SVs from whole-genome sequencing data remains challenging.
- Existing SV genotypers struggle with variant and sample-specific biases.
Purpose of the Study:
- To introduce NPSV, a novel machine learning approach for genotyping known structural variants.
- To address biases in next-generation sequencing data for improved SV genotyping.
- To enhance the accuracy of structural variant detection in genomic analyses.
Main Methods:
- Developed NPSV, a machine learning-based genotyper.
- Utilized next-generation sequencing simulation to model genomic, sequencer, and alignment pipeline effects.
- Evaluated NPSV against existing genotypers on benchmark call sets.
Main Results:
- NPSV achieves state-of-the-art genotyping accuracy across diverse SV call sets, samples, and variant types.
- Demonstrated consistent or superior performance compared to existing methods.
- Successfully identified de novo SVs in trio contexts and showed robustness to breakpoint offsets.
Conclusions:
- Stand-alone SV genotyping is increasingly vital with growing SV databases and long-read sequencing data.
- NPSV offers a robust framework for accurate SV genotyping in various applications by simulating biases.
- The approach facilitates accurate genotyping of a wide range of SVs in targeted and genome-scale studies.
More Related Videos
Related Concept Videos
Genome-wide Association Studies-GWAS
14.8K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
14.8K
Next-generation Sequencing
95.0K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
95.0K
Genomics
38.3K
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
38.3K
Comparing Copy Number Variations and SNPs
18.2K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
18.2K
Genome Annotation and Assembly
19.7K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
19.7K
Sanger Sequencing
763.4K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
763.4K

