Related Experiment Video
Updated: Sep 8, 2025

09:45
Detection of Copy Number Alterations Using Single Cell Sequencing
Published on: February 17, 2017
11.7K
Polishing copy number variant calls on exome sequencing data via deep learning
Furkan Özden1, Can Alkan1, A Ercüment Çiçek1,2
1Department of Computer Engineering, Bilkent University, 06800 Ankara, Turkey.
Genome Research
|June 13, 2022
Summary
A new deep learning model, DECoNT, enhances copy number variant (CNV) detection using whole-exome sequencing (WES) data. This improves precision for duplication and deletion calls, making WES a more reliable tool for genetic disease research.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Copy number variants (CNVs) are crucial in genetic diseases.
- Whole-genome sequencing (WGS) offers high accuracy for CNV detection.
- Whole-exome sequencing (WES) is cost-effective but less accurate for CNVs due to capture biases.
Purpose of the Study:
- To develop a deep learning model to improve CNV detection accuracy on WES data.
- To correct inaccuracies in existing WES-based CNV callers.
- To enhance the reliability of germline CNV detection using ubiquitous WES data.
Main Methods:
- Developed DECoNT, a novel deep learning model.
- Trained DECoNT using matched WES and WGS data from the 1000 Genomes Project.
- Evaluated DECoNT's performance across different sequencing technologies, capture kits, and CNV callers.
Main Results:
- DECoNT significantly enhances CNV detection accuracy on WES data.
- Achieved a threefold increase in duplication call precision and a twofold increase in deletion call precision.
- Demonstrated consistent performance improvement irrespective of sequencing or capture methods.
Conclusions:
- DECoNT acts as a universal polishing tool for exome CNV calls.
- The model substantially improves the reliability of germline CNV detection from WES data.
- Facilitates more accurate genetic disease association studies using cost-efficient WES data.
Related Concept Videos
Comparing Copy Number Variations and SNPs
17.9K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.9K
Genome Copying Errors
4.4K
DNA replication is a well-evolved process that copies millions of base pairs with high fidelity during each cell division. Occasionally a wrong base or a long stretch of wrong bases may get added to the daughter strands. If the errors are left unchecked, cells might accumulate several mutations that might endanger their survival. Therefore, the copying errors are checked and repaired at three levels.
4.4K
Next-generation Sequencing
92.5K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
92.5K
Improving Translational Accuracy
11.8K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.8K
Sanger Sequencing
756.9K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
756.9K

