Related Experiment Video
Updated: Dec 15, 2025

Rare Event Detection Using Error-corrected DNA and RNA Sequencing
Published on: August 3, 2018
REINDEER: efficient indexing of k-mer presence and abundance in sequencing datasets
Camille Marchet1, Zamin Iqbal2, Daniel Gautheret3
1CNRS, UMR 9189 - CRIStAL, Université de Lille, F-59000 Lille, France.
Motivation:
In this work we present REINDEER, a novel computational method that performs indexing of sequences and records their abundances across a collection of datasets. To the best of our knowledge, other indexing methods have so far been unable to record abundances efficiently across large datasets.
Results:
We used REINDEER to index the abundances of sequences within 2585 human RNA-seq experiments in 45 h using only 56 GB of RAM. This makes REINDEER the first method able to record abundances at the scale of ∼4 billion distinct k-mers across 2585 datasets. REINDEER also supports exact presence/absence queries of k-mers. Briefly, REINDEER constructs the compacted de Bruijn graph of each dataset, then conceptually merges those de Bruijn graphs into a single global one. Then, REINDEER constructs and indexes monotigs, which in a nutshell are groups of k-mers of similar abundances.
Availability And Implementation:
https://github.com/kamimrcht/REINDEER.
Supplementary Information:
Supplementary data are available at Bioinformatics online.
Related Concept Videos
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Multi-species Conserved Sequences
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Maxam-Gilbert Sequencing
Challenges of the Maxam-Gilbert Method
The...
Sanger Sequencing
Tagging and Fusion Proteins
Modern Molecular Taxonomy

