Related Experiment Video
Updated: Nov 29, 2025

CIRCLE-Seq for Interrogation of Off-Target Gene Editing
Published on: November 1, 2024
CRISPR sequences are sometimes erroneously translated and can contaminate public databases with spurious proteins
Alejandro Rubio1, Pablo Mier2, Miguel A Andrade-Navarro2
1Centro Andaluz de Biologia del Desarrollo (CABD, UPO-CSIC-JA). Facultad de Ciencias Experimentales (Área de Genética), Universidad Pablo de Olavide, Ctra. Utrera, Km.1, 41013, Sevilla, Spain.
Abstract:
The genomics era is resulting in the generation of a plethora of biological sequences that are usually stored in public databases. There are many computational tools that facilitate the annotation of these sequences, but sometimes they produce mistakes that enter the databases and can be propagated when erroneous data are used for secondary analyses, such as gene prediction or homology searching. While developing a computational gene finder based on protein-coding sequences, we discovered that the reference UniProtKB protein database is contaminated with some spurious sequences translated from DNA containing clustered regularly interspaced short palindromic repeats. We therefore encourage developers of prokaryotic computational gene finders and protein database curators to consider this source of error.
Related Concept Videos
Leaky Scanning
CRISPR
CRISPR/Cas9 Genome Editing
Conservative Site-specific Recombination and Phase Variation
The recognition sites for Cre recombinase called LoxP...
The Antiviral System of Bacteria and Archaea: CRISPR
CRISPR and crRNAs
The CRISPR-Cas system stores a copy of foreign DNA in the host genome and uses it to identify the foreign DNA upon reinfection. CRISPR-Cas has three different...

