Related Experiment Video
Updated: Jan 3, 2026

09:45
Detection of Copy Number Alterations Using Single Cell Sequencing
Published on: February 17, 2017
12.0K
Improving Copy Number Variant Detection from Sequencing Data with a Combination of Programs and a Predictive Model.
Salla Välipakka1, Marco Savarese1, Lydia Sagath1
1Folkhälsan Research Center, Helsinki, Finland.
The Journal of Molecular Diagnostics : JMD
|November 17, 2019
Summary
This study introduces a bioinformatics pipeline for detecting copy number variants (CNVs) using gene panel massively parallel sequencing (MPS) data. Combining four CNV detection tools and a statistical model improves accuracy for neuromuscular disorder research.
Area of Science:
- Genomics
- Bioinformatics
- Genetic Disorders
Background:
- Copy number variants (CNVs) analysis from gene panel massively parallel sequencing (MPS) data is challenging due to less developed bioinformatics tools.
- Accurate CNV detection is crucial for diagnosing genetic disorders, including neuromuscular disorders.
Purpose of the Study:
- To develop and validate an efficient bioinformatics pipeline for CNV detection from gene panel MPS data.
- To evaluate the performance of different CNV detection programs and their combinations.
- To establish a statistical model for improving the accuracy and standardization of CNV detection.
Main Methods:
- In silico generation of CNVs in MPS gene panel data.
- Analysis of in silico CNVs using four complementary programs: CoNIFER, XHMM, ExomeDepth, and CODEX.
- Training and validation of a logistic regression model using CNV detection scores from the four programs.
Main Results:
- A combination of all four CNV detection programs yielded more sensitive results than individual programs or other combinations.
- The logistic regression model incorporating scores from all four programs demonstrated superior performance in validation.
- No single program achieved sufficient accuracy for detecting all CNV sizes and types.
Conclusions:
- A combination of carefully selected bioinformatics tools is essential for maximizing CNV detection accuracy.
- A statistical model aids in streamlining and standardizing the filtering and annotation of detected CNVs.
- The developed pipeline offers an efficient approach for CNV analysis in gene panel MPS data for neuromuscular disorders.
Related Concept Videos
Comparing Copy Number Variations and SNPs
18.5K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
18.5K
Genome Copying Errors
5.0K
DNA replication is a well-evolved process that copies millions of base pairs with high fidelity during each cell division. Occasionally a wrong base or a long stretch of wrong bases may get added to the daughter strands. If the errors are left unchecked, cells might accumulate several mutations that might endanger their survival. Therefore, the copying errors are checked and repaired at three levels.
5.0K
Next-generation Sequencing
97.4K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
97.4K
RNA-seq
11.6K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
11.6K
Sanger Sequencing
772.5K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
772.5K

