使用标签数组进行无损的pangenome索引
Parsa Eskandar1, Benedict Paten1, Jouni Sirén1
1UC Santa Cruz Genomics Institute, University of California, Santa Cruz, 1156 High Street, Santa Cruz, CA 95064, USA.
bioRxiv : the preprint server for biology
|June 4, 2025
概括
我们开发了一个新的索引框架,使用标签数组来高效地查询复杂的 pangenome 图形. 这种方法可以在大型基因组中进行无损的,对哈普洛型有意识的搜索,从而改进了基因组分析工具.
科学领域:
- 生物信息学是一种生物信息学.
- 计算型基因组学计算型基因组学
- 数据结构 数据结构
背景情况:
- 泛基因组图对于在多个单基因组型中表示基因组变异至关重要.
- 大规模泛基因组数据的高效和无损索引仍然是一个重要的计算挑战.
研究的目的:
- 为了呈现一个实用和可扩展的索引框架,为 pangenome 图表.
- 为了使复杂的泛基因组结构中的查询模式能够有效和无损地检索.
主要方法:
- 开发了一个标签数组索引框架,将FM-index扩展到图形坐标.
- 介绍了一种结合k-mers,图形扩展和单元型穿越的新型构造算法.
- 利用多串的布罗斯-惠勒转换 (BWT) 和r-index来进行大基因组的存储效率高的处理.
主要成果:
- 标签数组结构证明了有效的压缩和可扩展性,随着随机类型的增加.
- 精确的映射信息在不同的基因组区域中被保存.
- 该方法可以在复杂的基因组中进行无损且对哈普罗型有意识的查询.
结论:
- 拟议的索引框架为查询复杂的基因组提供了一个实用的解决方案.
- 这种方法促进了可扩展的对齐器和基于图形的生物信息学工具的开发.
更多相关视频
07:09Author Spotlight: Navigating Challenges and Innovations in Muscle Stem Cell Studies
Published on: March 1, 2024
888
12:29Identifying Transcription Factor Olig2 Genomic Binding Sites in Acutely Purified PDGFRα+ Cells by Low-cell Chromatin Immunoprecipitation Sequencing Analysis
Published on: April 16, 2018
9.4K
相关概念视频
Tagging and Fusion Proteins
7.1K
Proteins are involved in several cellular processes and biochemical reactions. Analyzing a specific protein of interest requires it to be isolated from the other proteins in the cell. This is achieved by overexpressing the specific gene in a suitable host to produce large quantities of the target protein. A tag or label is recombined with the gene to produce a fusion protein containing the target protein and the tag. The tags on these fusion proteins can then be used for easy detection and...
7.1K
Genome Annotation and Assembly
19.4K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
19.4K
DNA Microarrays
18.7K
Microarrays are high-throughput and relatively inexpensive assays that can be automated to analyze large quantities of data at a time. They are used in genome-wide studies to compare gene or protein expression under two varied conditions, such as healthy and diseased states. Microarrays consist of glass or silica slides on which probe molecules are covalently attached through surface functionalization. Most commonly, the slides are prepared through the chemisorption of silanes to silica...
18.7K
Multi-species Conserved Sequences
4.3K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
4.3K
Genomic DNA in Eukaryotes
47.8K
Eukaryotes have large genomes compared to prokaryotes. To fit their genomes into a cell, eukaryotic DNA is packaged extraordinarily tightly inside the nucleus. To achieve this, DNA is tightly wound around proteins called histones, which are packaged into nucleosomes that are joined by linker DNA and coil into chromatin fibers. Additional fibrous proteins further compact the chromatin, which is recognizable as chromosomes during certain phases of cell division.
47.8K
Labeling DNA Probes
8.4K
DNA probes are fragments of DNA labeled with a reporter tag to enable their detection or purification. The resulting labeled DNA probes can then hybridize to target nucleic acid sequences through complementary base-pairing, and may be used to recover or identify these regions.
Radioisotopes, fluorophores, or small molecule binding partners like biotin or digoxigenin, are the most widely used reporter tags for labeling DNA probes. These labels can be attached to the probe DNA molecule via...
Radioisotopes, fluorophores, or small molecule binding partners like biotin or digoxigenin, are the most widely used reporter tags for labeling DNA probes. These labels can be attached to the probe DNA molecule via...
8.4K
