Related Experiment Video
Updated: Jan 15, 2026

09:40
Novel Sequence Discovery by Subtractive Genomics
Published on: January 25, 2019
9.1K
Efficient and accurate search in petabase-scale sequence repositories
Mikhail Karasikov1,2,3, Harun Mustafa1,2,3, Daniel Danciu1,3
1Biomedical Informatics Group, Department of Computer Science, ETH Zurich, Zurich, Switzerland.
Nature
|October 8, 2025
Summary
We developed MetaGraph, a framework to make vast biological sequence data full-text searchable. This cost-effective solution enables efficient searching of millions of DNA, RNA, and protein sequences for biomedical research.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Public biological sequencing data is rapidly expanding, posing challenges for efficient full-text searching.
- Existing methods struggle to handle the scale and complexity of large sequence datasets.
Purpose of the Study:
- To present MetaGraph, a scalable framework for indexing and full-text searching of biological sequence data.
- To demonstrate the cost-effectiveness and feasibility of searching massive sequence repositories.
Main Methods:
- Utilizing annotated de Bruijn graphs for scalable indexing of DNA, RNA, and protein sequences.
- Integrating data from seven public sources, encompassing 18.8 million unique sequence sets and 210 billion amino acid residues.
- Developing a cost-effective search mechanism for large sequence repositories.
Main Results:
- Achieved full-text searchability for 18.8 million unique DNA/RNA sequence sets and 210 billion amino acid residues.
- Demonstrated cost-effective searching of 67 petabase pairs of raw sequence data, with costs as low as $0.74 per queried megabase pair.
- Showcased that a compressed representation of all public biological sequences can fit on consumer hard drives for approximately $2,500.
Conclusions:
- MetaGraph provides a cost-effective and scalable solution for full-text searching of biological sequence data.
- The framework facilitates integrative analyses and has the potential to accelerate biomedical research advancements.
- Accessible and searchable large-scale biological data can catalyze new discoveries and applications.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
6.8K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
6.8K
Multi-species Conserved Sequences
4.6K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
4.6K
Protein Families
16.7K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
16.7K

