Related Experiment Video
Updated: Jul 13, 2025

10:36
Rare Event Detection Using Error-corrected DNA and RNA Sequencing
Published on: August 3, 2018
12.1K
Evaluation of haplotype-aware long-read error correction with hifieval.
Yujie Guo1,2, Xiaowen Feng1, Heng Li1,2
1Department of Data Science, Dana-Farber Cancer Institute, Boston, MA 02215, United States.
Bioinformatics (Oxford, England)
|October 18, 2023
Summary
A new tool, hifieval, evaluates error correction in PacBio High-Fidelity (HiFi) sequencing data. This assessment helps improve the accuracy of de novo sequence assemblers for better genomic analysis.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- PacBio High-Fidelity (HiFi) sequencing generates highly accurate long reads (>99%).
- De novo sequence assemblers utilize error correction (EC) as a primary step for processing HiFi data.
- The performance of EC algorithms on this new data type remains unevaluated.
Purpose of the Study:
- To introduce hifieval, a novel command-line tool for quantifying over- and under-corrections in EC algorithms.
- To assess the accuracy of EC components within existing HiFi assemblers.
- To investigate EC performance in challenging genomic regions.
Main Methods:
- Development of the hifieval tool for measuring EC accuracy.
- Evaluation of EC components from current HiFi assemblers.
- Testing on CHM13 and HG002 datasets, including analysis of homopolymer, centromeric, and segmental duplication regions.
Main Results:
- The study provides the first evaluation of EC accuracy for HiFi sequencing data.
- Performance variations in EC methods were identified across different genomic contexts.
- hifieval enables quantitative assessment of EC over- and under-corrections.
Conclusions:
- hifieval is a valuable tool for benchmarking and improving EC algorithms.
- Enhanced EC accuracy is crucial for advancing de novo assembly quality with HiFi data.
- This work will guide the development of more robust HiFi assemblers.
Related Concept Videos
Genome Copying Errors
4.2K
DNA replication is a well-evolved process that copies millions of base pairs with high fidelity during each cell division. Occasionally a wrong base or a long stretch of wrong bases may get added to the daughter strands. If the errors are left unchecked, cells might accumulate several mutations that might endanger their survival. Therefore, the copying errors are checked and repaired at three levels.
4.2K
Long-patch Base Excision Repair
7.0K
Since the discovery of the two BER pathways, there has been a debate about how a cell chooses one pathway over the other and the factors determining this selection. Numerous in vitro experiments have pointed out multiple determinants for the sub-pathway selection. These are:
7.0K
Mismatch Repair
40.2K
Overview
40.2K
Improving Translational Accuracy
11.4K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.4K
Homologous Recombination
50.6K
The basic reaction of homologous recombination (HR) involves two chromatids that contain DNA sequences sharing a significant stretch of identity. One of these sequences uses a strand from another as a template to synthesize DNA in an enzyme-catalyzed reaction. The final product is a novel amalgamation of the two substrates. To ensure an accurate recombination of sequences, HR is restricted to the S and G2 phases of the cell cycle. At these stages, the DNA has been replicated already and the...
50.6K
Sanger Sequencing
754.6K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
754.6K

