Troubleshooting common errors in assemblies of long-read metagenomes.
Florian Trigodet1,2, Rohan Sachdeva3, Jillian F Banfield4,5,6,7,8
1Helmholtz Institute for Functional Marine Biodiversity, Oldenburg, Germany.
Nature Biotechnology
|January 2, 2026
Summary
Evaluating long-read metagenome assemblies reveals significant errors, including chimeras and repeat issues, impacting genome recovery accuracy. A new tool aids in assessing these assembly errors for improved reliability.
Area of Science:
- Genomics
- Bioinformatics
- Metagenomics
Background:
- Assessing long-read assembly accuracy in complex environmental metagenomes is difficult.
- Underrepresented organisms pose additional challenges for genome recovery.
Purpose of the Study:
- To benchmark four leading long-read assembly software programs.
- To identify and quantify errors in metagenome assemblies.
- To develop a reproducible workflow for assembly evaluation.
Main Methods:
- Benchmarking HiCanu, hifiasm-meta, metaFlye, and metaMDBG on 21 PacBio HiFi metagenomes.
- Quantifying read clipping events during mapping to assembled contigs.
- Analyzing mock communities, gut microbiomes, and ocean samples.
Main Results:
- Long-read metagenome assemblies can contain over 40 errors per 100 million base pairs.
- Identified errors include multi-domain chimeras, premature circularization, haplotyping errors, excessive repeats, and phantom sequences.
- Significant divergence between assembled contigs and source reads was observed.
Conclusions:
- Current long-read assembly methods exhibit substantial error rates in metagenomic contexts.
- A novel open-source tool and workflow facilitate rigorous evaluation of assembly errors.
- Improved genome recovery from metagenomes requires addressing these identified assembly inaccuracies.
Related Concept Videos
Genome Annotation and Assembly
20.5K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
20.5K
Genome Copying Errors
5.0K
DNA replication is a well-evolved process that copies millions of base pairs with high fidelity during each cell division. Occasionally a wrong base or a long stretch of wrong bases may get added to the daughter strands. If the errors are left unchecked, cells might accumulate several mutations that might endanger their survival. Therefore, the copying errors are checked and repaired at three levels.
5.0K
Mismatch Repair
43.4K
Overview
43.4K


