Related Experiment Video
Updated: Dec 29, 2025

09:45
Detection of Copy Number Alterations Using Single Cell Sequencing
Published on: February 17, 2017
12.0K
Rapid, Paralog-Sensitive CNV Analysis of 2457 Human Genomes Using QuicK-mer2
Feichen Shen1, Jeffrey M Kidd2
1Department of Human Genetics, University of Michigan, Ann Arbor, MI 48109, USA.
Genes
|February 5, 2020
Summary
QuicK-mer2 is a new tool that rapidly maps gene copy-number variation without genome mapping. This approach reveals extensive genetic diversity in paralogous genes across large populations.
Area of Science:
- Genomics and Evolutionary Biology
- Population Genetics
- Bioinformatics
Background:
- Gene duplication drives the evolution of novel gene functions and contributes significantly to genetic diversity through copy-number variation (CNV).
- Existing CNV detection methods often struggle with duplicated sequences and are computationally intensive, hindering analysis of large population datasets.
- Distinguishing between highly similar paralogous genes is a key challenge in accurately assessing CNV.
Purpose of the Study:
- To develop a rapid, mapping-free computational approach for constructing paralog-specific copy-number maps from short-read sequencing data.
- To enable efficient analysis of copy-number variation in paralogous genes across large populations.
- To identify novel instances of copy-number variation in human populations.
Main Methods:
- Developed QuicK-mer2, a self-contained, mapping-free bioinformatics tool.
- Utilizes tabulation of unique k-mer sequences from short-read data to identify paralog-specific copy numbers.
- Applied QuicK-mer2 to 2457 individuals from the 1000 Genomes Project dataset, analyzing a 20X human genome in approximately 20 minutes.
Main Results:
- Successfully constructed paralog-specific copy-number maps for 2457 individuals.
- Identified widespread copy-number variation in paralogous genes, including nine genes with no samples having a copy number of two.
- Discovered 92 genes with a majority of samples exhibiting a copy number other than two, and characterized rare CNV at the APOBEC3 locus.
Conclusions:
- QuicK-mer2 provides a fast and computationally efficient method for analyzing paralog-specific copy-number variation.
- The study reveals substantial uncharacterized copy-number variation in human paralogous genes.
- This approach facilitates the study of genetic diversity and evolutionary dynamics of gene families in population-scale genomics.
More Related Videos
Related Concept Videos
Comparing Copy Number Variations and SNPs
18.5K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
18.5K
Genome Copying Errors
4.9K
DNA replication is a well-evolved process that copies millions of base pairs with high fidelity during each cell division. Occasionally a wrong base or a long stretch of wrong bases may get added to the daughter strands. If the errors are left unchecked, cells might accumulate several mutations that might endanger their survival. Therefore, the copying errors are checked and repaired at three levels.
4.9K
Single Nucleotide Polymorphisms-SNPs
17.8K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
17.8K
Gene Duplication and Divergence
7.7K
The seminal work of Ohno in 1970 popularized the idea of gene duplication and divergence. DNA sequence comparison studies reveal that a large portion of the genes in bacteria, archaebacteria, and eukaryotes was generated by gene duplication and divergence, indicating its critical role in evolution.
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are...
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are...
7.7K

