Related Experiment Video
Updated: May 16, 2026

09:37
An Integrated Approach for Microprotein Identification and Sequence Analysis
Published on: July 12, 2022
Scaffolding low quality genomes using orthologous protein sequences
1Wellcome Trust Centre for Human Genetics, University of Oxford, Oxford, OX3 7BN, UK. yang.li@well.ox.ac.uk
Bioinformatics (Oxford, England)
|November 20, 2012
Summary
SWiPS (Scaffolding With Protein Sequences) improves fragmented genome assemblies using orthologous proteins. This method enhances gene space representation and scaffold accuracy, even with low-quality data.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Next-generation sequencing generates highly fragmented genome assemblies.
- Low-quality assemblies pose challenges for downstream genomic analyses.
- Existing assembly methods struggle with large intergenic regions in eukaryotes.
Purpose of the Study:
- To present SWiPS (Scaffolding With Protein Sequences), a novel pipeline for improving fragmented genome assemblies.
- To utilize orthologous protein sequences as guides for scaffolding and gene prediction.
- To enhance the quality and completeness of low-quality genome assemblies.
Main Methods:
- SWiPS employs orthologous protein sequences to scaffold existing contigs.
- The pipeline simultaneously predicts gene structures through homology.
- It does not require high N50 values or complete proteins on single contigs.
Main Results:
- SWiPS improved N50 by ~20% for *Ciona intestinalis* and *Homo sapiens* assemblies.
- Gene space representation increased by >110% for *Callorhinchus milii* and 20-40% for *C. intestinalis*.
- Scaffold error rates were low, with 85-90% of scaffolds fully correct and >95% of local contig joins accurate.
Conclusions:
- SWiPS effectively improves fragmented genome assemblies using protein sequence information.
- The method enhances gene space recovery and scaffold accuracy in diverse eukaryotic genomes.
- SWiPS offers a valuable tool for researchers working with low-quality or fragmentary genome assemblies.
More Related Videos
Related Concept Videos
Genome Annotation and Assembly
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
Evolutionary Relationships through Genome Comparisons
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
Gene Families
Gene families consist of groups of genes proposed to have originated from a common ancestor. Typically these arise through events in which a gene or genes are mistakenly duplicated during cell division. Unlike their parent genes (which are subject to selection pressure to maintain function), these gene copies do not need to preserve their sequences and may evolve at a relatively faster rate.
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
Protein Families
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key locations, protein...
Conservation of Protein Domains Over Different Proteins
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to form...
Protein Complexes with Interchangeable Parts
Groups of proteins may form a complex where each protein in this complex has a different role in the overall execution of the complex’s function. Often some of the proteins in the complex can be replaced by a closely related variant to give a complex that contains many of the same components yet is functionally distinct.
The SCF ubiquitin ligase is a protein complex of five individual proteins. This complex attaches ubiquitin to other target proteins to mark them for degradation. In order to...
The SCF ubiquitin ligase is a protein complex of five individual proteins. This complex attaches ubiquitin to other target proteins to mark them for degradation. In order to...

