タンパク質ファミリーにおける機能的関係を明らかにするための構造認識型生成AIフレームワーク
Divyanshu Shukla1, Jonathan Martin2, Faruck Morcos2,3,4
1Bioinformatics and Computational Biology Program, Iowa State University.
bioRxiv : the preprint server for biology
|December 25, 2025
まとめ
本研究では、3Dジオメトリと共進化を利用してタンパク質配列を解析する新しいフレームワークを紹介する。構造情報を統合することで、相同性検出とタンパク質設計を改善する。
科学分野:
- 計算生物学
- 構造バイオインフォマティクス
- タンパク質科学
背景:
- タンパク質配列データベースは、実験的な構造決定よりも速く成長しており、多くの配列が注釈なしで残っている。
- タンパク質フォールドは配列よりも保存性が高く情報量が多い、強力な解析ツールを提供する。
研究 の 目的:
- タンパク質配列のマッピング、クラスタリング、設計のための生成的な構造認識型フレームワークを開発すること。
- 幾何学的エンコーディングと共進化制約を活用して、タンパク質解析を強化すること。
主な方法:
- タンパク質の局所構造の離散幾何学的表現のための3D相互作用(3Di)アルファベットを採用した。
- アミノ酸配列と3Di表現間の双方向翻訳のためにProstT5を利用した。
- 潜在的生成ランドスケープ内に3Diアラインメントを直接カップリング解析(DCA)および変分オートエンコーダー(VAE)と統合した。
主要な成果:
- 高感度な相同性検出と構造誘導型タンパク質配列生成を達成した。
- 共進化シグナルの検出を強化し、構造変異体を合理的にサンプリングした。
- 多様なタンパク質ファミリーにわたる接触予測、相同性推論、配列生成の改善を実証した。
結論:
- 統合フレームワークは、タンパク質構造空間の定量的かつ生成的なビューを提供する。
- 構造情報を取り入れることで、タンパク質の進化と設計の研究を前進させる。
- 低い配列同一性を持つ遠縁の相同性であっても、高感度な解析と設計を可能にする。
関連する概念動画
Protein Families
16.6K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
16.6K
Protein Families
4.1K
4.1K
Gene Families
3.5K
3.5K
Gene Families
9.7K
Gene families consist of groups of genes proposed to have originated from a common ancestor. Typically these arise through events in which a gene or genes are mistakenly duplicated during cell division. Unlike their parent genes (which are subject to selection pressure to maintain function), these gene copies do not need to preserve their sequences and may evolve at a relatively faster rate.
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
Occasionally these regions can be adapted to take on new roles within the organism, becoming novel genes...
9.7K
Conservation of Protein Domains Over Different Proteins
14.0K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
14.0K
Protein Networks
4.4K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.4K


