在伪随机和真实序列中的同义和停止编码子的统计分析作为GC内容的函数
Valentin Wesp1, Günter Theißen2, Stefan Schuster3
1Department of Bioinformatics, Matthias Schleiden Institute, Friedrich Schiller University Jena, Ernst-Abbe-Platz 2, 07743, Jena, Germany.
Scientific reports
|December 27, 2023
概括
DNA中的同义符号频率受到GC含量的影响,影响基因发现和开放的读取框架识别. 这项研究分析了25个遗传代码中的这些频率,揭示了信息内容的最佳GC内容.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 分子生物学分子生物学
背景情况:
- 在DNA中同义三重频率对于发现基因和理解基因组组成至关重要.
- GC内容显著影响这些频率,特别是在停止编码子和开放读取框架.
- 之前的研究已经探讨了代码使用偏差,但需要对多个遗传代码进行全面的分析,作为GC内容的函数.
研究的目的:
- 根据它们的编码组计算所有25个遗传密码的氨基酸和停止编码频率.
- 分析这些频率作为伪随机DNA序列中GC含量的函数.
- 确定基于同名编码子的Shannon信息最大化的GC内容.
主要方法:
- 利用伪随机序列模型和所有25个已知的遗传代码进行分析.
- 计算了氨基酸和停止编码子频率作为GC含量的函数.
- 根据同名的编码组确定了Shannon信息内容.
主要成果:
- 氨基酸根据其在特定GC含量水平的最大频率被分为五组.
- 标准遗传密码的最大Shannon信息含量发生在43.3%的GC含量.
- 自然序列表现出一种偏见,即在各种类别的内部中,在各种类别的内部中,在拼接位附近的停止编码子.
结论:
- GC 含量是 DNA 中同义密码子频率和氨基酸组成的关键决定因素.
- 观察到的信息含量的最佳GC含量与大多数真核生物的基因组参数保持一致.
- 在拼接部位附近停止密码子分布表明内子的功能或进化意义.
相关概念视频
Gene Evolution - Fast or Slow?
7.1K
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
7.1K
From DNA to Protein
18.4K
The flow of genetic information in cells from DNA to mRNA to protein is described by the central dogma, which states that genes specify the sequence of mRNAs, which in turn specify the sequence of amino acids making up all proteins. The decoding of one molecule to another is performed by specific proteins and RNAs. Because the information stored in DNA is so central to cellular function, it makes intuitive sense that the cell would make mRNA copies of this information for protein synthesis...
18.4K
Multi-species Conserved Sequences
3.9K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
3.9K
Single Nucleotide Polymorphisms-SNPs
15.1K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
15.1K
Comparing Copy Number Variations and SNPs
17.7K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.7K
Evolutionary Relationships through Genome Comparisons
5.7K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.7K


