HVRLocator: a computationally efficient tool for identifying hypervariable regions in large 16S rRNA datasets
Clara Arboleda-Baena1,2, Felipe Borim Correa2, Joao Pedro Saraiva2
1German Centre for Integrative Biodiversity Research (iDiv) Halle-Jena-Leipzig, Puschstraße 4, 04103 Leipzig, Germany.
Gigascience
|April 11, 2026
Summary
HVRLocator accurately identifies 16S rRNA gene regions and primers in metabarcoding data. This tool enhances microbial diversity studies by improving data curation and enabling reliable cross-study comparisons.
Area of Science:
- Microbiology
- Bioinformatics
- Genomics
Background:
- 16S rRNA gene metabarcoding is a cost-effective method for microbial diversity assessment.
- Public datasets often lack standardized metadata, hindering accurate data reuse.
- Critical metadata includes sequenced hypervariable regions and primers used.
Purpose of the Study:
- Introduce HVRLocator, a computational tool for analyzing 16S rRNA metabarcoding data.
- Address the challenge of missing or unreliable metadata in public sequence archives.
- Enable accurate curation and large-scale processing of 16S rRNA data.
Main Methods:
- HVRLocator identifies 16S rRNA amplicon start/end positions.
- The tool determines corresponding hypervariable regions.
- It detects the presence of primer sequences within the data.
Main Results:
- HVRLocator processes archived sequences at 6.5 samples/minute.
- It accurately detects amplicon positions and hypervariable regions across diverse datasets.
- The tool flags misannotated metadata and problematic sequences, aiding data curation.
Conclusions:
- HVRLocator overcomes metadata limitations in 16S rRNA studies.
- Accurate identification of amplicon regions and primers ensures reliable data processing.
- Enables reproducible microbial studies, syntheses, and meta-analyses.
Related Concept Videos
Gene Evolution - Fast or Slow?
8.4K
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
8.4K
Cis-regulatory Sequences
12.3K
Cis-regulatory sequences are short fragments of non-coding DNA that are present on the same chromosomes as the genes that they regulate. These fragments serve as binding sites for transcriptional regulators, proteins that are responsible for controlling gene transcription and differential gene expression across cell types in eukaryotes. Cis-regulatory sequences can be close to the gene of interest or thousands of bases away in the DNA sequence; however, those sequences that are further away are...
12.3K
Cis-regulatory Sequences
4.4K
4.4K
Sanger Sequencing
780.1K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
780.1K
Genomic DNA in Eukaryotes
54.3K
Eukaryotes have large genomes compared to prokaryotes. To fit their genomes into a cell, eukaryotic DNA is packaged extraordinarily tightly inside the nucleus. To achieve this, DNA is tightly wound around proteins called histones, which are packaged into nucleosomes that are joined by linker DNA and coil into chromatin fibers. Additional fibrous proteins further compact the chromatin, which is recognizable as chromosomes during certain phases of cell division.
54.3K
RNA-seq
12.6K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
12.6K


