蛋白质结构预测需要多少元基因组数据:从生态和进化角度来看,有针对性的方法的优点
1Key Laboratory of Molecular Biophysics of the Ministry of Education, Hubei Key Laboratory of Bioinformatics and Molecular-Imaging, Department of Bioinformatics and Systems Biology Center of AI Biology, College of Life Science and Technology, Huazhong University of Science and Technology Wuhan Hubei China.
iMeta
|June 13, 2024
概括
超基因组数据有助于蛋白质结构建模,但许多蛋白质仍然未解决. 本综述确定了元基因组数据中的生态和进化模式,以改善蛋白质结构预测,突出了针对性的方法,以获得更好的结果.
科学领域:
- 基因组学就是基因组学.
- 结构生物学 结构生物学
- 生物信息学是一种生物信息学.
背景情况:
- 可以使用同源和元基因组序列建模三维蛋白质结构.
- 尽管有大量的元基因组数据,但大量的蛋白质结构仍然未解决.
- 了解元基因组数据中的生态和进化模式对于改善蛋白质结构预测至关重要.
研究的目的:
- 在元基因组数据中识别生态和进化模式.
- 解码这些模式和蛋白质结构之间的关系.
- 调查这些模式的有效使用,以提高蛋白质结构预测.
主要方法:
- 提出了一种元基因组利用效率和边际效应模型来量化同源序列分布.
- 对比了针对性与非针对性的方法来识别来自特定生物群的同源序列.
- 确定了预测所有Pfam数据库蛋白质结构所需的元基因组数据的下界.
主要成果:
- 发现了与蛋白质结构预测相关的元基因组数据中的生态和进化模式.
- 有针对性的方法有效地识别特定生物组的同源序列.
- 目前的元基因组数据不足以预测Pfam数据库中的所有蛋白质结构.
结论:
- 可以利用元基因组数据中的生态和进化模式进行有效的蛋白质结构预测.
- 有针对性的方法显示出提取同源序列和改善蛋白质结构预测的前景.
- 需要进一步采集数据和改进方法来解决目前的数据不足问题.
相关概念视频
Evolutionary Relationships through Genome Comparisons
5.7K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.7K
Proteomics
7.3K
A proteome is the entire set of proteins that a cell type produces. We can study proteomes using the knowledge of genomes because genes code for mRNAs, and the mRNAs encode proteins. Although mRNA analysis is a step in the right direction, not all mRNAs are translated into proteins.
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term...
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term...
7.3K
Genome Annotation and Assembly
18.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.8K
Genomics
36.3K
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
36.3K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Conservation of Protein Domains Over Different Proteins
10.8K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.8K


