Related Experiment Video
Updated: May 26, 2026

11:04
RNA Next-Generation Sequencing and a Bioinformatics Pipeline to Identify Expressed LINE-1s at the Locus-Specific Level
Published on: May 19, 2019
LOCALE: Local-Alignment Embeddings for Noise-Robust DNA Search at SRA Scale
Ryan P Synk1, Prashant Pandey2, S Cenk Sahinalp3
1University of Maryland, College Park.
Biorxiv : the Preprint Server for Biology
|May 25, 2026
Summary
We developed LOCALE, a new method for searching large sequencing datasets. LOCALE uses vector embeddings to accurately find related DNA sequences, even with errors or mutations.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Searching large-scale raw sequencing data repositories like the NIH Sequence Read Archive (SRA) is crucial for biological discovery.
- Current search methods struggle with scalability and are sensitive to sequencing errors and biological variations due to reliance on exact k-mer matching.
Purpose of the Study:
- To develop a scalable and robust sequence search method for petabase-scale repositories.
- To improve the accuracy of sequence retrieval in the presence of sequencing errors and biological divergence.
Main Methods:
- Recasting sequence search as a dense retrieval problem using vector embeddings.
- Training a DNABERT-2 encoder with an InfoNCE objective on biologically informed data augmentations (corrupted sequence crops).
- Evaluating the LOCALE method on SRA benchmarks with varying dataset sizes and mutation rates.
Main Results:
- LOCALE achieved 62.4% average Recall@Rq at a 10% mutation rate on a 50-accession SRA benchmark, outperforming baselines in noisy-query settings.
- On a larger 500-accession, 15-Gbp benchmark, LOCALE achieved an AUPRC of 0.508 at 10% mutation, significantly higher than MetaGraph's 0.129.
- The method demonstrates effective retrieval by ranking locally aligned sequences higher than unaligned ones.
Conclusions:
- LOCALE offers a scalable and accurate solution for searching large sequencing data repositories.
- The dense retrieval approach effectively handles sequencing errors and biological divergence, outperforming traditional methods.
- This method has the potential to transform biological discovery by enabling efficient exploration of vast genomic datasets.
Related Concept Videos
DNA Microarrays
Microarrays are high-throughput and relatively inexpensive assays that can be automated to analyze large quantities of data at a time. They are used in genome-wide studies to compare gene or protein expression under two varied conditions, such as healthy and diseased states. Microarrays consist of glass or silica slides on which probe molecules are covalently attached through surface functionalization. Most commonly, the slides are prepared through the chemisorption of silanes to silica...
RNA-seq
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while microarray-based...
Conserved Binding Sites
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally analyses the...
Nucleic Acid Structure
The pentose sugar in DNA is deoxyribose, while in RNA the pentose sugar is ribose. The difference between the sugars is the presence of the hydroxyl group on the ribose's second carbon and a hydrogen on the deoxyribose's second carbon. The phosphate residue attaches to the hydroxyl group of the 5′ carbon of one sugar and the hydroxyl group of the 3′ carbon of the sugar of the next nucleotide, which forms a 5′ to 3′ phosphodiester linkage.
DNA Structure
DNA has a double-helix structure. The...
DNA Structure
DNA has a double-helix structure. The...
Base-pairing and DNA Repair
Erwin Chargaff’s rules on DNA equivalence paved the way for the discovery of base pairing in DNA. Chargaff’s rules state that in a double-stranded DNA molecule,
DNA Base Pairing
Erwin Chargaff’s rules on DNA equivalence paved the way for the discovery of base pairing in DNA. Chargaff’s rules state that in a double-stranded DNA molecule,

