Related Experiment Video
Updated: Sep 29, 2025

12:08
Hybrid De Novo Genome Assembly for the Generation of Complete Genomes of Urinary Bacteria using Short- and Long-read Sequencing Technologies
Published on: August 20, 2021
5.2K
Haplotype-resolved assembly of diploid genomes without parental data
Haoyu Cheng1,2, Erich D Jarvis3,4, Olivier Fedrigo3
1Department of Data Science, Dana-Farber Cancer Institute, Boston, MA, USA.
Nature Biotechnology
|March 25, 2022
Summary
Generating haplotype-resolved genome assemblies from single samples is now possible. Our new algorithm uses PacBio HiFi and Hi-C data to create high-quality, phased genome assemblies without parental DNA.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Haplotype-resolved genome assembly is crucial for understanding genetic variation.
- Current methods often require parental samples, limiting their applicability.
Purpose of the Study:
- To develop a novel algorithm for haplotype-resolved genome assembly from single samples.
- To eliminate the need for parental DNA in generating phased genome assemblies.
Main Methods:
- Integration of PacBio HiFi long reads with Hi-C chromatin interaction data.
- Development of a computational pipeline to process and assemble the combined data.
Main Results:
- The algorithm successfully produced haplotype-resolved genome assemblies for human and vertebrate samples.
- The method consistently outperformed existing single-sample assembly pipelines.
- Assembly quality was comparable to pedigree-based methods.
Conclusions:
- This algorithm provides a robust solution for single-sample haplotype-resolved genome assembly.
- It significantly advances the field by removing the dependency on parental sequencing.
- Enables broader application of phased genome assemblies in research and clinical settings.
Related Concept Videos
Genome Annotation and Assembly
19.4K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
19.4K
Evolutionary Relationships through Genome Comparisons
6.3K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
6.3K
Polytene Chromosomes
10.3K
Polytene chromosomes are giant interphase chromosomes with several DNA strands placed side by side. They were discovered in the year 1881 by Balbiani in salivary glands, intestine, muscles, malpighian tubules, and hypoderm of larvae Chironomus plumosus. Hence, these are also called "Salivary gland chromosomes." These are found in insects of the order Diptera and Collembola; in certain organs of mammals; and synergids, antipodes of flowering plants. Polytene chromosomes are also...
10.3K
Genome-wide Association Studies-GWAS
14.5K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
14.5K

