Related Experiment Video
Updated: Jun 28, 2025

Mapping Bacterial Functional Networks and Pathways in Escherichia Coli using Synthetic Genetic Arrays
Published on: November 12, 2012
Many purported pseudogenes in bacterial genomes are bona fide genes
Nicholas P Cooley1, Erik S Wright2,3
1Department of Biomedical Informatics, University of Pittsburgh, Pittsburgh, PA, USA.
Background:
Microbial genomes are largely comprised of protein coding sequences, yet some genomes contain many pseudogenes caused by frameshifts or internal stop codons. These pseudogenes are believed to result from gene degradation during evolution but could also be technical artifacts of genome sequencing or assembly.
Results:
Using a combination of observational and experimental data, we show that many putative pseudogenes are attributable to errors that are incorporated into genomes during assembly. Within 126,564 publicly available genomes, we observed that nearly identical genomes often substantially differed in pseudogene counts. Causal inference implicated assembler, sequencing platform, and coverage as likely causative factors. Reassembly of genomes from raw reads confirmed that each variable affects the number of putative pseudogenes in an assembly. Furthermore, simulated sequencing reads corroborated our observations that the quality and quantity of raw data can significantly impact the number of pseudogenes in an assembler dependent fashion. The number of unexpected pseudogenes due to internal stops was highly correlated (R2 = 0.96) with average nucleotide identity to the ground truth genome, implying relative pseudogene counts can be used as a proxy for overall assembly correctness. Applying our method to assemblies in RefSeq resulted in rejection of 3.6% of assemblies due to significantly elevated pseudogene counts. Reassembly from real reads obtained from high coverage genomes showed considerable variability in spurious pseudogenes beyond that observed with simulated reads, reinforcing the finding that high coverage is necessary to mitigate assembly errors.
Conclusions:
Collectively, these results demonstrate that many pseudogenes in microbial genome assemblies are actually genes. Our results suggest that high read coverage is required for correct assembly and indicate an inflated number of pseudogenes due to internal stops is indicative of poor overall assembly quality.
More Related Videos
08:34Generation of In-Frame Gene Deletion Mutants in Pseudomonas aeruginosa and Testing for Virulence Attenuation in a Simple Mouse Model of Infection
Published on: January 8, 2020
12:29Generation of Null Mutants to Elucidate the Role of Bacterial Glycosyltransferases in Bacterial Motility
Published on: March 11, 2022
Related Concept Videos
Genome Size and the Evolution of New Genes
Genomic DNA in Prokaryotes
Genomic Diversity in Bacteria
Although bacterial genomes are much...
Antibiotic Selection
Gene Duplication and Divergence
The duplicated copies of the gene are called Paralogs. Paralogs with similar sequences and functions form a gene family. Across several species, a large number of gene families are...
Reporter Genes