Related Experiment Video
Updated: Jun 4, 2025

Heuristic Mining of Hierarchical Genotypes and Accessory Genome Loci in Bacterial Populations
Published on: December 7, 2021
Evaluating data requirements for high-quality haplotype-resolved genomes for creating robust pangenome references
Prasad Sarashetti1, Josipa Lipovac2, Filip Tomas2
1Laboratory of Human Genomics, Genome Institute of Singapore, A*STAR, Singapore, Singapore.
This study provides guidance on optimal data types and volumes for robust de novo genome assembly in population-level pangenome projects. PacBio HiFi and ONT Duplex reads, along with ultra-long and long-range data, are crucial for high-quality phased genomes.
Area of Science:
- Genomics
- Computational Biology
- Population Genetics
Background:
- Long-read sequencing technologies (PacBio HiFi, Duplex, ultra-long ONT) enable advanced genome assembly.
- Pangenome references are crucial for representing genetic diversity but face challenges in data selection and cost.
- Lack of clear guidance hinders optimal data choices for pangenome studies.
Purpose of the Study:
- To evaluate optimal data types and volumes for de novo genome assembly in population-level pangenome projects.
- To compare the performance of Oxford Nanopore Technologies (ONT) Duplex and Pacific Biosciences (PacBio) HiFi datasets for phased genome assembly.
- To provide recommendations for cost-effective and sensitive pangenome assembly.
Main Methods:
- Comparative analysis of PacBio HiFi and ONT Duplex long-read sequencing data.
- Assessment of ultra-long ONT reads and long-range data (Omni-C, Hi-C) for genome assembly.
- Evaluation of contiguity, completeness, and phasing accuracy of resulting genome assemblies.
Main Results:
- Chromosome-level haplotype-resolved assembly requires approximately 20× high-quality long reads (HiFi or Duplex) per haplotype.
- 15–20× ultra-long ONT reads and 10× long-range data are recommended for robust assembly.
- Both HiFi and Duplex yield comparable contiguity; HiFi excels in phasing accuracy, while Duplex produces more Telomere-to-Telomere (T2T) contigs.
Conclusions:
- The study offers insights into optimal data selection for population-level pangenome projects.
- Reassessing recommended data types and volumes considering economic constraints is vital for the pangenome research community.
- These findings will aid in advancing genomic studies with broader impacts.
More Related Videos
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Genome Annotation and Assembly
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Gene Evolution - Fast or Slow?
In contrast, regions which code...

