CAZyme3D:碳水化合物活性酶的3D结构数据库
bioRxiv : the preprint server for biology
|January 7, 2025
概括
CAZyme3D是一个新的碳水化合物活性酶 (CAZymes) 3D结构数据库. 它使用结构相似性来分类酶,有助于发现各种应用的新型CAZymes.
科学领域:
- 生物化学和结构生物学
- 生物信息学和计算生物学
背景情况:
- 碳水化合物活性酶 (CAZymes) 对包括人类健康,营养和生物能源在内的生物过程至关重要.
- 目前的CAZyme注释方法仅依赖于序列相似性,限制了发现进化上遥远但功能上相关的酶.
- 在CAZymes的专用3D结构数据库中存在一个缺口,这阻碍了基于结构相似性的分析.
研究的目的:
- 开发CAZyme3D,为CAZymes提供一个全面的3D结构数据库.
- 建立基于结构相似性的CAZymes的新型层次分类系统.
- 为结构相似性搜索提供一个工具,以帮助发现新的CAZymes.
主要方法:
- 使用AlphaFold生成了870,740个3D结构,形成了CAZyme3D整个数据集.
- 对188,574个非冗余的CAZyme序列 (ID50数据集) 进行了基于结构相似性的聚类.
- 开发了一个等级分类,包括现有的CAZy级别和新的子类,结构集群 (SC) 组和SC级别.
主要成果:
- 跨家族聚类将具有相似结构折叠的CAZy家族和氏族分组为子类.
- 内部家族聚类确定了结构相似的CAZymes,分为SCs和SC组,与基于序列的子家族不同.
- CAZyme3D提供了一个基于Web的平台,用于提交查询序列或PDB结构以进行结构相似性搜索.
结论:
- 通过利用3D结构信息,CAZyme3D为CAZyme注释和发现提供了一种强大的新方法.
- 结构分类系统增强了对超越序列同质性的CAZyme关系的理解.
- 预计这项资源将加速在各种研究领域识别和描述新型CAZymes.
相关概念视频
Gene Families
8.7K
Gene families consist of groups of genes proposed to have originated from a common ancestor. Typically these arise through events in which a gene or genes are mistakenly duplicated during cell division. Unlike their parent genes (which are subject to selection pressure to maintain function), these gene copies do not need to preserve their sequences and may evolve at a relatively faster rate.
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
8.7K
Protein Organization
6.2K
Proteins are polymers of amino acid residues. They are versatile and responsible for different cellular functions, including DNA replication, molecular transport, catalysis, and structural support. Proteins have a hierarchical structure comprising at least three levels of organization: primary, secondary, and tertiary structure. Some large proteins have a quaternary structure where individual protein subunits are linked together.
The primary structure of a protein is its amino acid sequence....
The primary structure of a protein is its amino acid sequence....
6.2K


