TetRex:一种用于指数加速搜索高度保存图案的新算法
Remy M Schwab1,2, Simon Gene Gottlieb2, Knut Reinert1,2
1Max-Planck-Institute for Molecular Genetics, Ihnestrasse 63, 14195 Berlin, Germany.
NAR genomics and bioinformatics
|April 18, 2025
概括
本研究引入了一种用于在大型生物数据集中搜索正则表达式 (regex) 的新算法. 通过使用一种新的数据结构,它显著加快了对保存的遗传序列的搜索过程.
科学领域:
- 计算生物学 计算生物学
- 生物信息学是一种生物信息学.
- 基因组学就是基因组学.
背景情况:
- 现代数据集需要先进的计算解决方案来解决基本问题,如集合成员和核心值.
- 跨物种保存的生物序列通常使用正则表达式 (regexes) 来表示.
- 当前的regex搜索工具与生物序列数据库的规模扎.
研究的目的:
- 在大型生物序列数据库中开发一种新的算法,用于高效的regex搜索.
- 为了提高扫描巨大的基因组和蛋白质组数据集的性能.
主要方法:
- 开发了一个新的regex搜索算法.
- 集成了一个层次交叉的布鲁姆过器,为集成成员提供了一个紧的数据结构.
- 结合了新的算法与传统的搜索方法.
主要成果:
- 新的算法有效地减少了对regex查询的搜索空间.
- 与现有的最先进的工具相比,这种综合方法显示出更高的性能.
- 实现了更快的运行时间来扫描大型生物序列数据库.
结论:
- 拟议的算法为生物序列分析提供了重大进步.
- 利用像Bloom过器这样的紧数据结构可以优化计算密集型任务.
- 这项创新解决了现代生物数据规模所带来的挑战.
相关概念视频
Conserved Binding Sites
4.1K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.1K
Multi-species Conserved Sequences
3.8K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
3.8K
Conservation of Protein Domains
3.0K
3.0K
Conservation of Protein Domains Over Different Proteins
10.6K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.6K
Cis-regulatory Sequences
2.9K
2.9K


