SnoBIRD:一种工具,用于识别C/D盒 snoRNA,并在所有真核生物中改进它们的注释
Étienne Fafard-Couture1,2, Cédric Boulanger1,2, Laurence Faucher-Giguère2,3
1Département de biochimie et de génomique fonctionnelle, Faculté de médecine et des sciences de la santé, Université de Sherbrooke, Sherbrooke, Québec J1E 4K8, Canada.
Nucleic acids research
|July 30, 2025
概括
我们开发了SnoBIRD,这是一种用于识别幼核RNA (snoRNA) 和它们在真核生物基因组中的伪基因的新工具. SnoBIRD准确地预测C/D盒 snoRNAs,改善了基因组注释和进化研究.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 分子生物学分子生物学
背景情况:
- 小核RNAs (snoRNAs) 是关键的非编码RNAs,参与核糖体生物发生和跨真核细胞的拼接.
- 现有的snoRNAs基因组注释往往不完整,无法区分功能性snoRNA和伪基因.
研究的目的:
- 开发一个准确和高效的计算工具,用于预测C/D盒子snoRNA及其在真核生物基因组中的伪基因.
- 提高snoRNA注释的完整性和准确性,并促进进化分析.
主要方法:
- 开发了基于BERT的预测器SnoBIRD,该预测器在多种类型的真核生物C/D盒 snoRNA中进行训练.
- 使用生物相关信号对比SnoBIRD与现有的预测工具.
- 在裂变酵母,人类和多个真核生物基因组中应用SnoBIRD.
主要成果:
- 与现有的工具相比,SnoBIRD在预测snoRNA和识别snoRNA伪基因方面表现出卓越的性能.
- 该工具与基因组大小具有出色的可扩展性,在运行时间效率方面表现优于其他预测器.
- 斯诺BIRD成功识别了许多新的C/D盒 snoRNA 和伪基因,包括在裂变酵母和人类基因组中经过实验验证的候选人.
- 跨多个真核生物基因组的分析揭示了数百个新的snoRNA候选者,并提供了对snoRNA演变的见解.
结论:
- 斯诺BIRD是一个用户友好,高效和可靠的工具,用于在任何真核生物基因组中全面的C / D盒斯诺RNA和伪基因预测.
- 该工具显著提高了基因组注释的准确性,并有助于理解snoRNAs的进化动态.
相关概念视频
Comparing Copy Number Variations and SNPs
17.9K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.9K
Genome Annotation and Assembly
19.3K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
19.3K
Genomic DNA in Eukaryotes
47.7K
Eukaryotes have large genomes compared to prokaryotes. To fit their genomes into a cell, eukaryotic DNA is packaged extraordinarily tightly inside the nucleus. To achieve this, DNA is tightly wound around proteins called histones, which are packaged into nucleosomes that are joined by linker DNA and coil into chromatin fibers. Additional fibrous proteins further compact the chromatin, which is recognizable as chromosomes during certain phases of cell division.
47.7K


