Related Experiment Video
Updated: Nov 17, 2025

10:36
Rare Event Detection Using Error-corrected DNA and RNA Sequencing
Published on: August 3, 2018
12.3K
Factorial estimating assembly base errors using k-mer abundance difference (KAD) between short reads and genome
Cheng He1, Guifang Lin1, Hairong Wei2
1Department of Plant Pathology, Kansas State University, 4024 Throckmorton Center, Manhattan, KS 66506-5502, USA.
NAR Genomics and Bioinformatics
|February 12, 2021
Summary
A new method called k-mer abundance difference (KAD) assesses genome assembly quality by comparing k-mer copy numbers. KAD identifies base errors, insertions, deletions, and redundancy, aiding precise error correction.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Genome sequencing and assembly are crucial for understanding genetic content.
- Long-read sequencing has improved genome contiguity but introduces errors.
- Existing error correction methods struggle with residual errors in assemblies.
Purpose of the Study:
- To develop a novel approach for evaluating genome assembly quality.
- To identify and quantify various types of errors in assembled genome sequences.
- To provide a diagnostic tool for precise error correction in genome assemblies.
Main Methods:
- Developed the k-mer abundance difference (KAD) method.
- Compared inferred k-mer copy numbers from short reads with observed copy numbers in the assembly.
- Utilized KAD metrics to classify k-mers and assess assembly quality.
Main Results:
- KAD successfully identifies base errors and estimates overall error rates in genome assemblies.
- The method detects sequence insertions, deletions, and redundancy.
- KAD metrics provide insights into assembly quality and error profiles.
Conclusions:
- KAD is a valuable tool for the quality evaluation of genome assemblies.
- KAD facilitates the identification of specific error types, aiding targeted correction.
- The developed KAD software is available for public use to improve genome assembly accuracy.
Related Concept Videos
Genome Annotation and Assembly
19.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
19.8K
Genome Copying Errors
4.8K
DNA replication is a well-evolved process that copies millions of base pairs with high fidelity during each cell division. Occasionally a wrong base or a long stretch of wrong bases may get added to the daughter strands. If the errors are left unchecked, cells might accumulate several mutations that might endanger their survival. Therefore, the copying errors are checked and repaired at three levels.
4.8K

