基因模型的保存可以支持基因组注释
Cassandria G Tay Fernandez1, Philipp E Bayer1, Jakob Petereit1
1School of Biological Sciences and Institute of Agriculture, University of Western Australia, Perth, Western Australia, Australia.
The plant genome
|August 21, 2023
概括
基因组注释中的假阳性基因模型可能会导致错误. 本研究介绍了一种使用进化保护的方法来识别和纠正这些错误,为豆类基因组注释提供了宝贵的资源.
科学领域:
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
- 进化生物学 进化生物学
背景情况:
- 基因组注释通常包含错误阳性基因模型.
- 这些错误会影响遗传学和比较基因组分析.
- 准确的基因模型对于理解基因组功能和进化至关重要.
研究的目的:
- 开发一种方法来支持使用进化保护的基因模型预测.
- 在豆类基因组中识别潜在错误的基因注释.
- 为豆类研究创建一套高质量的代表性基因模型.
主要方法:
- 利用进化保护模式来评估基因模型的有效性.
- 开发了一种用于基因模型预测和验证的计算方法.
- 将该方法应用于12种豆类基因组组合.
主要成果:
- 成功识别了可能错误的基因注释.
- 创建了一个15345个代表性豆类基因模型的数据集.
- 提出的方法在改善基因组注释质量方面表现出有效性.
结论:
- 进化保护是验证基因模型的可靠指标.
- 开发的方法和基因模型集可以提高豆类基因组注释的准确性.
- 这项工作为豆类中更可靠的比较基因组学提供了基础.
相关概念视频
Genome Annotation and Assembly
18.9K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.9K
Conservation of Protein Domains
3.1K
3.1K
Conservation of Protein Domains Over Different Proteins
10.9K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.9K
Multi-species Conserved Sequences
4.0K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
4.0K
DNA as a Genetic Template
22.0K
Two structural features of the DNA molecule provide a basis for the mechanisms of heredity: the four nucleotide bases and its double-stranded nature. The Watson-Crick model of double-helical DNA structure, proposed in 1952, drew heavily upon the X-ray crystallography work of researchers Rosalind Franklin and Maurice Wilkins. Watson, Crick, and Wilkins jointly received the Nobel Prize in Physiology or Medicine for their work in 1962. Franklin was, controversially, excluded from the prize for...
22.0K
Genome Size and the Evolution of New Genes
8.0K
While every living organism has a genome of some kind (be it RNA, or DNA), there is considerable variation in the sizes of these blueprints. One major factor that impacts genome size is whether the organism is prokaryotic or eukaryotic. In prokaryotes, the genome contains little to no non-coding sequence, such that genes are tightly clustered in groups or operons sequentially along the chromosome. Conversely, the genes in eukaryotes are punctuated by long stretches of non-coding sequence.
8.0K


