高质量的证据,以注释非正规的开放阅读框架作为人类蛋白质
Eric W Deutsch1, Leron W Kok2,3, Jonathan M Mudge4
1Institute for Systems Biology (ISB), Seattle, WA, 98109, USA.
bioRxiv : the preprint server for biology
|September 24, 2024
概括
研究人员已经确定了超过25%的非正典开放阅读框架 (ncORF) 的翻译产品,大大扩大了我们对人类蛋白质组及其在健康和疾病中的作用的理解.
科学领域:
- 基因组学和蛋白质组学
- 人类健康和疾病
背景情况:
- 编码蛋白质的基因组对于研究人类健康至关重要,但之前的分析可能错过了重要的组成部分.
- 在人类细胞和疾病中观察到非正典的开放阅读框架 (ncORF),这对蛋白质组学,基因组学和临床科学有潜在的影响.
- 对ncORF对人类蛋白质的贡献缺乏大规模的理解,从而限制了它们的影响.
研究的目的:
- 通过合作努力,建立一个关于ncORFs的蛋白质水平证据的共识格局.
- 通过一组定义的ncORF来确定翻译的程度,并对产生的进行表征.
- 为编写ncORFs制定一个框架,并为研究人员创建可访问的公共工具.
主要方法:
- 来自蛋白质组学,免疫组学,Ribo-seq ORF发现和基因注释的综合数据.
- 在泛蛋白质组分析中分析了95520个实验中的38亿个质谱.
- 开发了ncORF研究的注释框架和公共工具 (GENCODE,PeptideAtlas).
主要成果:
- 在7,264个调查的ncORF中,至少有25%的蛋白质水平证据被发现.
- 从翻译的ncORF中鉴定出超过3000种独特的.
- 为ncORFs建立了一个全面的数据集和注释框架.
结论:
- 这项工作通过表征ncORF衍生的蛋白质,大大提高了对人类蛋白质的理解.
- 开发的资源为未来利用ncORF数据进行生物医学发现提供了一个平台.
- 这些发现对人类健康研究有意义,并扩展到ncORF对其他生物的观察.
关键词:
在GENCODE中,我们可以使用GENCODE.人类蛋白质组项目里博-seqq 的意思免疫类药物 免疫类药物质谱测量质谱测量质谱测量质量测量质谱测量质量测量质量测量质量测量质量测量质量测量质量测量质量测量质量测量质量测量质量测量质量测量质量测量质量测量质量测量质量测量质量测量质量测量质量测量质量测量质量测量微型蛋白质是一种微型蛋白质.非正规的ORFs是非正规的.蛋白质组学 蛋白质组学翻译翻译翻译翻译翻译翻译更多相关视频
07:38Mass Spectrometry-Based Proteomics Analyses Using the OpenProt Database to Unveil Novel Proteins Translated from Non-Canonical Open Reading Frames
Published on: April 11, 2019
12.7K
08:23De novo Identification of Actively Translated Open Reading Frames with Ribosome Profiling Data
Published on: February 18, 2022
3.5K
相关概念视频
Leaky Scanning
5.1K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.1K
Ribosome Profiling
3.5K
Ribosome profiling or ribo-sequencing is a deep sequencing technique that produces a snapshot of active translation in a cell. It selectively sequences the mRNAs protected by ribosomes to get an insight into a cell’s translation landscape at any given point in time.
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique...
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique...
3.5K
Signal Sequences and Sorting Receptors
5.3K
Signal sequences are short amino acid sequences that guide newly synthesized proteins to their proper location within the cell. Classical signal sequences are fifteen to sixty amino acids long and present at the N-terminus of a polypeptide chain. Each signal sequence has a conserved segment of basic residues towards their N terminus, a hydrophobic core, and a C-terminus rich in polar residues. The C-terminus also contains a signal cleavage site and features a -3 -1 sequence motif. The -3-1...
5.3K
Proteins: From Genes to Degradation
12.1K
Within a biological system, the DNA encodes the RNA, and the nucleotide sequence in the RNA further defines the amino acid sequence in the protein. This is referred to as “The Central Dogma of Molecular Biology” - a term coined by Francis Crick. Central dogma is a firm principle in biology that defines the flow of genetic information within any life form. The two fundamental steps in central dogma are - transcription and translation.
Transcription is the synthesis of RNA...
Transcription is the synthesis of RNA...
12.1K
Genome Annotation and Assembly
18.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.8K
Conservation of Protein Domains Over Different Proteins
10.8K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.8K
