GenDiS3数据库:在整个序列数据库中对已知结构的蛋白质域超级家族在整个序列数据库中的流行率进行普查
Sarthak Joshi1, Shailendu Mohapatra2, Dhwani Kumar1
1National Centre for Biological Sciences, Tata Institute of Fundamental Research, GKVK Campus, Bellary Road, Bangalore 560065, India.
概括
基因DiS3数据库使用先进的生物信息学将蛋白质序列与已知的结构联系起来,识别了超过15100万个同类. 本资源有助于理解蛋白质的演变和功能.
科学领域:
- 结构生物学 结构生物学
- 生物信息学是一种生物信息学.
- 基因组学就是基因组学.
背景情况:
- 在现有的蛋白质序列数据和实验确定的3D结构之间存在很大的差距.
- 弥合这种序列结构差距对于理解蛋白质功能和进化至关重要.
- 计算工具对于分类蛋白序列和预测结构至关重要.
研究的目的:
- 更新和增强GenDiS数据库 (超级家族的基因组分布) 以改善蛋白质序列结构链接.
- 为识别和验证蛋白质域超级家族同类物提供一个强大的资源.
- 为了促进蛋白质的功能注释和进化研究.
主要方法:
- 利用先进的生物信息工具,包括DELTA-BLAST用于初始同类体检测和HMMSCAN用于验证.
- 更新了GenDiS数据库,包含了蛋白质域超级家族的基因组分布信息.
- 对庞大的序列数据库进行大规模的计算搜索.
主要成果:
- 在2060个超级家族 (SCOPe) 中识别了超过15100万个序列同类.
- 验证了1.16亿个这些同类的真实阳性,显著提高了准确性.
- 对糖解酶和LOG基因的案例研究揭示了进化变化和功能多样性.
结论:
- 更新的GenDiS3数据库为研究人员提供了强大的资源.
- GenDiS3通过将序列与已知的结构联系起来,有助于功能注释和进化研究.
- 数据库和相关工具可以在https://caps.ncbs.res.in/gendis3/.in/访问.
相关概念视频
Protein Families
15.2K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.2K
Conservation of Protein Domains Over Different Proteins
10.6K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.6K
Conservation of Protein Domains
3.0K
3.0K
Gene Families
8.7K
Gene families consist of groups of genes proposed to have originated from a common ancestor. Typically these arise through events in which a gene or genes are mistakenly duplicated during cell division. Unlike their parent genes (which are subject to selection pressure to maintain function), these gene copies do not need to preserve their sequences and may evolve at a relatively faster rate.
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
8.7K
Conserved Binding Sites
4.1K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.1K
Multi-species Conserved Sequences
3.9K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
3.9K


