GSR-DB:用于16S rRNA amplicon分析的手动策划和优化的分类学数据库
Leidy-Alejandra G Molano1, Sara Vega-Abellaneda1, Chaysavanh Manichanh1,2
1Microbiome Lab, Vall d'Hebron Institut de Recerca (VHIR), Vall d'Hebron Barcelona Hospital Campus, Passeig Vall d'Hebron, Barcelona, Spain.
mSystems
|January 8, 2024
概括
新的Greengenes,SILVA和RDP数据库 (GSR-DB) 改进了从16S amplicon测序的微生物分类学赋值. 与现有数据库相比,GSR-DB提供了增强的物种级分辨率,有助于微生物组研究.
科学领域:
- 微生物学 微生物学
- 生物信息学是一种生物信息学.
- 基因组学就是基因组学.
背景情况:
- 基于Amplicon的16S核糖体RNA测序对于微生物社区的分析至关重要,特别是在低生物质样本中.
- 现有的参考数据库 (例如,SILVA,Greengenes,GTDB,RDP) 存在不一致的命名和注释缺陷,限制了分类学分辨率.
- 这些局限性阻碍了在不同研究领域准确的微生物识别和分析.
研究的目的:
- 开发一个集成和手动策划的数据库,Greengenes,SILVA和RDP数据库 (GSR-DB),用于细菌和考古16S安普利康分类分析.
- 为了解决当前16S数据库中存在的命名和注释问题的不一致性.
- 为微生物分类学分配提供更准确,更全面的资源,增强物种级别的分辨率.
主要方法:
- 来自Greengenes,SILVA和RDP数据库的序列的集成和手动策划.
- 实施一个独特的分类统一步骤,以确保整合数据库中的注释一致.
- 使用模拟社区,真实数据集和十倍交叉验证的验证,将GSR-DB与现有数据库进行比较 (Greengenes,Greengenes 2,GTDB,ITGDB,SILVA,RDP,MetaSquare).
主要成果:
- GSR-DB展示了16S序列的增强分类学注释,在模拟社区评估中在物种层面上优于现有的数据库.
- 通过十倍交叉验证的验证证实了GSR-DB的优越性能,除了Greengenes 2.
- 该数据库支持全长16S序列和常用的超变区 (V4,V1-V3,V3-V4,V3-V5).
结论:
- GSR-DB提供了一个强大的,准确的解决方案,用于微生物分类学分析,使用16S安普利康数据.
- 其增强的物种级别分辨率和一致的注释使其成为微生物组研究的宝贵工具.
- 对于计算资源有限的研究人员来说,GSR-DB是可访问的,促进了微生物生态学和相关领域的更广泛应用.
相关概念视频
RNA-seq
10.0K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
10.0K
Gene Evolution - Fast or Slow?
7.1K
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
7.1K


