Related Experiment Video
Updated: Aug 23, 2025

Genome-wide Surveillance of Transcription Errors in Eukaryotic Organisms
Published on: September 13, 2018
VeChat: correcting errors in long reads using variation graphs
Xiao Luo1,2, Xiongbin Kang1, Alexander Schönhuth3,4
1Genome Data Science, Faculty of Technology, Bielefeld University, Bielefeld, Germany.
Abstract:
Error correction is the canonical first step in long-read sequencing data analysis. Current self-correction methods, however, are affected by consensus sequence induced biases that mask true variants in haplotypes of lower frequency showing in mixed samples. Unlike consensus sequence templates, graph-based reference systems are not affected by such biases, so do not mistakenly mask true variants as errors. We present VeChat, as an approach to implement this idea: VeChat is based on variation graphs, as a popular type of data structure for pangenome reference systems. Extensive benchmarking experiments demonstrate that long reads corrected by VeChat contain 4 to 15 (Pacific Biosciences) and 1 to 10 times (Oxford Nanopore Technologies) less errors than when being corrected by state of the art approaches. Further, using VeChat prior to long-read assembly significantly improves the haplotype awareness of the assemblies. VeChat is an easy-to-use open-source tool and publicly available at https://github.com/HaploKit/vechat .
Related Concept Videos
Genome Copying Errors
Mismatch Repair
Comparing Copy Number Variations and SNPs
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Sanger Sequencing
Fixing Double-strand Breaks
Proofreading

