基因组语言模型 (gLMs) 解码细菌基因组,以改善基因预测和翻译启动地点识别
Genereux Akotenou1, Achraf El Allali1
1Bioinformatics Laboratory, College of Computing, University Mohammed VI Polytechnic, Lot 660, Hay Moulay Rachid, Ben Guerir 43150, Morocco.
Briefings in bioinformatics
|July 3, 2025
概括
基因组语言模型 (gLMs) 通过解释像人类语言一样的遗传序列,显著提高了细菌基因预测的准确性. 这种方法增强了编码序列识别和翻译启动站点预测,优于传统方法.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 准确的细菌基因预测对于理解微生物功能和生物技术至关重要.
- 传统的基因预测方法面临着复杂的遗传变异和新型序列的局限性.
- 基因组语言模型 (gLMs) 提供了一种灵感来自自然语言处理 (NLP) 的新方法来解释遗传信息.
研究的目的:
- 通过基因语言模型 (gLMs) 提高细菌基因预测的准确性.
- 开发一个两阶段的框架来识别编码序列 (CDS) 区域和翻译启动站点 (TIS).
- 评估gLM方法的性能与已建立的基因预测工具相比.
主要方法:
- 利用基于变压器的模型,特别是DNABERT,用于基因预测.
- 采用了两阶段的框架:CDS区域的识别,然后是TIS的改进.
- 在NCBI的精心策划的数据集上精心调整的DNABERT使用k-mer令牌化完成细菌基因组.
主要成果:
- 开发的gLM,GeneLM,在细菌基因预测准确度方面取得了显著的改善.
- 与Prodigal,GeneMark-HMM和Glimmer相比,GeneLM减少了错过的CDS预测,并增加了匹配的注释.
- 与传统方法相比,TIS预测在与实验验证的地点相比取得了更高的性能.
结论:
- 基因组语言模型 (gLMs) 在解码遗传信息方面显示出显著的潜力,用于精确的细菌基因组注释.
- 基因LM实现了最先进的性能,超过了传统的基因寻找器.
- 这一进步凸显了语言模型对基因组分析和生物技术的变革性影响.
相关概念视频
Ribosome Profiling
3.6K
Ribosome profiling or ribo-sequencing is a deep sequencing technique that produces a snapshot of active translation in a cell. It selectively sequences the mRNAs protected by ribosomes to get an insight into a cell’s translation landscape at any given point in time.
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique...
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique...
3.6K
Leaky Scanning
5.2K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.2K
Translation in Prokaryotes
204
Prokaryote translation is a complex, highly coordinated process that converts genetic information from mRNA into functional proteins. It involves three stages: initiation, elongation, and termination, each facilitated by specific molecular components.Initiation of TranslationThe process begins with the assembly of the ribosomal subunits and initiation factors on the mRNA. In bacteria, the 30S ribosomal subunit recognizes the Shine-Dalgarno sequence in the mRNA, a conserved region upstream of...
204
Initiation of Translation
34.6K
Initiating translation is complex because it involves multiple molecules. Initiator tRNA, ribosomal subunits, and eukaryotic initiation factors (eIFs) are all required to assemble on the initiation codon of mRNA. This process consists of several steps that are mediated by different eIFs.
First, the initiator tRNA must be selected from the pool of elongator tRNAs by eukaryotic initiation factor 2 (eIF2). The initiator tRNA (Met-tRNAi) has conserved sequence elements including modified bases at...
First, the initiator tRNA must be selected from the pool of elongator tRNAs by eukaryotic initiation factor 2 (eIF2). The initiator tRNA (Met-tRNAi) has conserved sequence elements including modified bases at...
34.6K
Modern Molecular Taxonomy
157
Advancements in molecular biology have revolutionized the identification and characterization of bacteria, with multiple methods leveraging DNA sequencing for enhanced precision. As sequencing technologies improve and costs decline, these approaches are increasingly used in clinical, environmental, and evolutionary studies.Multilocus Sequence Typing (MLST) examines several housekeeping genes, essential chromosomal genes encoding cellular functions, to distinguish strains. Approximately...
157
Genome Annotation and Assembly
19.3K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
19.3K


