使用大型语言模型和软对齐的病毒基因注释的改进
William L Harrigan1, Barbra D Ferrell2, K Eric Wommack2
1Hawai'i Institute of Marine Biology, University of Hawai'i at Mānoa, Honolulu, HI, 96822, USA.
BMC bioinformatics
|April 25, 2024
概括
大型语言模型 (LLM) 提供了一种新的方法来注释病毒蛋白序列. 这种基于嵌入的方法比传统方法提高了准确性和可解释性.
科学领域:
- 分子生物学分子生物学
- 生物信息学是一种生物信息学.
- 人工智能的人工智能
背景情况:
- 蛋白序列注释是分子生物学中的一个重大挑战,特别是在病毒蛋白中.
- 传统的同源性搜索方法 (对齐,k-mer,基于个人资料) 由于有限的同源性,与病毒序列作斗争.
- 大型语言模型 (LLM) 提出了一种使用嵌入的蛋白质序列注释的新方法.
研究的目的:
- 引入一种使用LLM嵌入的蛋白质序列注释的新方法.
- 为了解决传统方法在注释病毒蛋白中的局限性.
- 为了提高蛋白质注释的效率和可解释性.
主要方法:
- 开发一种软对齐算法,利用氨基酸嵌入相似性.
- 通过使用嵌入相似性绕过传统的分数矩阵.
- 为可解释性创建透明的,类似BLAST的对齐可视化.
主要成果:
- 软对齐算法在效率和可解释性方面超过了基于嵌入式的聚合模型.
- 该方法允许用户追踪同源氨基酸,并提供清晰的对齐可视化.
- 这种新的方法成功地注释了blastp和基于聚合的方法未能检测到的序列,如在病毒正统组和ViralZone数据库中验证的那样.
结论:
- 基于LLM的嵌入方法具有改善蛋白质序列注释的巨大潜力,特别是在病毒基因组学中.
- 这种方法为分子生物学中蛋白质功能推断提供了更有效和更准确的途径.
- 这些发现突出了将人工智能与传统生物研究相结合的有希望的进步,以增强注释.
相关概念视频
Genome Annotation and Assembly
18.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.8K
Leaky Scanning
5.1K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.1K
Improving Translational Accuracy
10.2K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
10.2K
RNA-seq
9.9K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
9.9K
Evolutionary Relationships through Genome Comparisons
5.7K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.7K
Viral Mutations
32.3K
A mutation is a change in the sequence of bases of DNA or RNA in a genome. Some mutations occur during replication of the genome due to errors made by the polymerase enzymes that replicate DNA or RNA. Unlike DNA polymerase, RNA polymerase is prone to errors because it is not capable of “proofreading” its work. Viruses with RNA-based genomes, like HIV, therefore accrue mutations faster than viruses with DNA-based genomes. Because mutation and recombination provide the raw material...
32.3K


