Related Experiment Video
Updated: Jul 31, 2025

High-throughput Physical Mapping of Chromosomes using Automated in situ Hybridization
Published on: June 28, 2012
Gaps and complex structurally variant loci in phased genome assemblies.
David Porubsky1, Mitchell R Vollger1, William T Harvey1
1Department of Genome Sciences, University of Washington School of Medicine, Seattle, Washington 98195, USA.
Human genome assembly still contains over 140 gaps, even with advanced methods. Most gaps and misorientations occur in repetitive DNA, impacting protein-coding genes and highlighting needs for better assembly algorithms.
Area of Science:
- Genomics
- Bioinformatics
- Human Genetics
Background:
- Phased human genome assembly has advanced using long-read data and parental or linked-read information.
- Despite progress, current methods like trio-based assembly (e.g., trio-hifiasm) still result in numerous gaps (over 140 per assembly).
Purpose of the Study:
- To analyze assembly gaps, breaks, and misorientations in a diverse set of human haploid assemblies.
- To compare phasing accuracy using Strand-seq versus parental data.
- To identify common locations and types of assembly gaps and assess their impact on protein-coding genes.
Main Methods:
- Analysis of 182 haploid assemblies from 77 human samples.
- Comparison of chromosome-wide phasing accuracy between trio-based methods and Strand-seq.
- Detailed examination of gap locations, focusing on repetitive elements like segmental duplications and satellite DNA.
- Estimation of DNA misorientations and identification of large-scale alignment discontinuities (deletions/insertions).
Main Results:
- Assembly gaps predominantly cluster in large, identical repeats (35.4% segmental duplications, 22.3% satellite DNA, 27.4% GA/AT-rich regions).
- 1513 protein-coding genes overlap assembly gaps, with 231 recurrently affected.
- 6-7 Mbp of DNA are misoriented per haplotype, with 81% corresponding to large inversion polymorphisms.
- Significant large-scale deletions (11.9 Mbp) and insertions (161.4 Mbp) per haploid genome identified, primarily in satellite DNA, but also in 230 euchromatic regions impacting 197 protein-coding genes.
Conclusions:
- Trio-based methods are the gold standard, but Strand-seq offers comparable chromosome-wide phasing accuracy.
- Repetitive regions, particularly segmental duplications and satellite DNA, are major sources of assembly gaps and variations.
- Incompletely assembled and variable regions, including those affecting protein-coding genes, are critical targets for improving genome assembly algorithms and pangenome representations.
More Related Videos
08:03Heuristic Mining of Hierarchical Genotypes and Accessory Genome Loci in Bacterial Populations
Published on: December 7, 2021
11:35Screening for Functional Non-coding Genetic Variants Using Electrophoretic Mobility Shift Assay EMSA and DNA-affinity Precipitation Assay DAPA
Published on: August 21, 2016
Related Concept Videos
Genome Annotation and Assembly
Gene Duplication and Divergence
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are...
Conservative Site-specific Recombination and Phase Variation
The recognition sites for Cre recombinase called LoxP...
Genome Copying Errors
Next-generation Sequencing
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Evolutionary Relationships through Genome Comparisons