D-sORF:与翻译机器相关的实验检测的小型开放式读取框架 (sORFs) 的准确ab initio分类
Nikos Perdikopanis1,2,3, Antonis Giannakakis4,5, Ioannis Kavakiotis3
1Department of Electrical and Computer Engineering, University of Thessaly, 38221 Volos, Greece.
Biology
|August 28, 2024
概括
D-sORF是一种机器学习工具,可以从基因组序列中准确识别小开放的读取框架 (sORF) 和它们的翻译. 这种计算方法有助于从非基因区域发现新的小蛋白质.
科学领域:
- 基因组学就是基因组学.
- 计算生物学 计算生物学
- 蛋白质组学是指蛋白质组学.
背景情况:
- 小开放的读取框架 (sORF) 在基因组中很普遍,越来越多的证据表明来自非基因区域的翻译.
- 来自sORFs的功能性已在各种各样的生物体中被确定.
- 准确的sORF注释对于高通量检测小蛋白质至关重要,但仍然是一个挑战.
研究的目的:
- 开发一个机器学习框架,用于准确预测和编码sORFs的分类.
- 提供一种具有成本效益的计算方法来检测sORF,以补充生物实验.
- 加强从非基因区域翻译的小蛋白质的识别.
主要方法:
- 介绍D-sORF,一个机器学习框架,集成核酸上下文和基因信息.
- D-sORF仅使用基因组序列来预测sORF的编码潜力,不包括保存参数.
- 在小型ORF上使用精度和准确度指标评估D-sORF的性能.
主要成果:
- 对于小型ORF,D-sORF实现了94.74%的精度和92.37%的准确性.
- 将D-sORF应用于与核糖体相关的sORF,在识别生成转录中,其性能与核糖体测序 (Ribo-Seq) 相似或优于其.
- 在识别非生产性核糖体结合事件时,D-sORF表现出较低的错误阳性率.
结论:
- D-sORF是一种有效的计算工具,用于识别编码sORF及其翻译.
- 该框架可以显著帮助过Ribo-Seq数据中的假阳性,提高小蛋白发现的可靠性.
- D-sORF促进了来自非基因区域的小蛋白质的高通量检测,推进了基因组和蛋白质组研究.
相关概念视频
Leaky Scanning
5.1K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.1K
Ribosome Profiling
3.5K
Ribosome profiling or ribo-sequencing is a deep sequencing technique that produces a snapshot of active translation in a cell. It selectively sequences the mRNAs protected by ribosomes to get an insight into a cell’s translation landscape at any given point in time.
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique...
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique...
3.5K
Initiation of Translation
31.8K
Initiating translation is complex because it involves multiple molecules. Initiator tRNA, ribosomal subunits, and eukaryotic initiation factors (eIFs) are all required to assemble on the initiation codon of mRNA. This process consists of several steps that are mediated by different eIFs.
First, the initiator tRNA must be selected from the pool of elongator tRNAs by eukaryotic initiation factor 2 (eIF2). The initiator tRNA (Met-tRNAi) has conserved sequence elements including modified bases at...
First, the initiator tRNA must be selected from the pool of elongator tRNAs by eukaryotic initiation factor 2 (eIF2). The initiator tRNA (Met-tRNAi) has conserved sequence elements including modified bases at...
31.8K
Termination of Translation
25.3K
The large ribosomal subunit has several important structures essential to translation. These include the peptidyl transferase center (PTC) - which is the site where the peptide bond is formed - and a large, internal, water-filled tube through which the nascent polypeptide moves. This latter structure is called the Peptide Exit Tunnel, and it begins at the PTC and spans the body of the large ribosomal subunit. During translation, as the nascent polypeptide chain is synthesized, it passes through...
25.3K
Directing Proteins to the Rough Endoplasmic Reticulum
7.2K
The organelle-specific signaling sequences direct proteins synthesized in the cytosol to their final destination like ER, mitochondria, peroxisomes, etc. Some of the proteins directed to ER are then trafficked via vesicles to other organelles within the cell or the extracellular environment through the Golgi complex. For example, the rough ER synthesizes soluble proteins for transportation to the lysosomes or secretion out of the cell. It can also synthesize transmembrane proteins that can...
7.2K
Improving Translational Accuracy
9.4K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
9.4K


