sOCP:基于TIS和框架内特征预测smORF编码潜力的框架,并在人类基因组中有效应用
Zhao Peng1, Jiaqiang Li2, Xingpeng Jiang2
1School of Life Sciences, and Hubei Key Laboratory of Genetic Regulation and Integrative Biology, Central China Normal University, Wuhan 430079, Hubei, People's Republic of China.
Briefings in bioinformatics
|April 11, 2024
概括
我们开发了sOCP,这是一个计算框架,用于预测具有编码潜力的小开放阅读框架 (smORF). 该工具有助于理解smORF在疾病和生物通路中的作用.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 分子生物学分子生物学
背景情况:
- 小型开放式阅读框架 (smORF) 越来越多地被认为是它们在关键的生物途径中的作用,影响糖尿病和癌症等疾病.
- 准确的smORFs的in silico预测对于分析大规模的omics数据至关重要.
研究的目的:
- 开发和验证一个新的计算框架,sOCP,用于预测smORFs的编码潜力.
- 识别人类基因组中的新型smORF并创建一个全面的数据库.
主要方法:
- 构建了人类smORFs的预测模型,使用框架内特征和在起始编码子附近的核酸偏差.
- 将sOCP模型应用于Rattus norvegicus,以评估跨物种性能.
- 扫描了人类基因组的smORF,包括那些具有非正规起始编码子的人,以建立一个数据库.
主要成果:
- 与现有方法相比,sOCP模型显示出更高的预测准确性,一个小的功能集可以防止过度装配.
- 该模型在人类和Rattus norvegicus中表现良好.
- 创建了大约100万个具有编码潜力的新型人类smORFs的数据库,其中72,000个位于lncRNA区域.
结论:
- sOCP框架为smORF预测提供了强大而有效的工具,增强了我们对其生物意义的理解.
- 确定了可能参与独特生物过程的新型smORFs,例如葡萄糖皮质类代谢和 prokaryotic 防御系统.
- 开发的数据库是未来对人类smORFs的研究的宝贵资源.
相关概念视频
Single Nucleotide Polymorphisms-SNPs
15.0K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
15.0K
Comparing Copy Number Variations and SNPs
17.7K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.7K
Cis-regulatory Sequences
9.9K
Cis-regulatory sequences are short fragments of non-coding DNA that are present on the same chromosomes as the genes that they regulate. These fragments serve as binding sites for transcriptional regulators, proteins that are responsible for controlling gene transcription and differential gene expression across cell types in eukaryotes. Cis-regulatory sequences can be close to the gene of interest or thousands of bases away in the DNA sequence; however, those sequences that are further away are...
9.9K
Ribosome Profiling
3.5K
Ribosome profiling or ribo-sequencing is a deep sequencing technique that produces a snapshot of active translation in a cell. It selectively sequences the mRNAs protected by ribosomes to get an insight into a cell’s translation landscape at any given point in time.
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique...
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique...
3.5K
Sanger Sequencing
754.2K
DNA sequencing is a fundamental technique that is routinely used in the biological sciences. This method can be applied to a range of questions at different scales - from the sequencing of a cloned DNA fragment or the study of a mutation in a gene up to whole-genome sequencing. However, despite the widespread use of sequencing today, it was not until 1977 that Fredrick Sanger and his collaborators developed the chain-termination method to decode DNA sequences. It relies on the separation of a...
754.2K
Improving Translational Accuracy
10.3K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
10.3K


