稀有的自编码器在蛋白质语言模型表示中发现了生物可解释的特征
Onkar Gujral1,2, Mihir Bafna2, Eric Alm3,4
1Department of Mathematics, Massachusetts Institute of Technology, Cambridge, MA 02139.
概括
稀疏的自编码器可以从蛋白质语言模型 (PLM) 中提取可解释的特征,而无需监督. 这些特征揭示了生物学见解,提高了AI在生命科学中的解释性和信任.
科学领域:
- 计算生物学 计算生物学
- 生命科学中的人工智能
- 生物信息学是一种生物信息学.
背景情况:
- 蛋白质语言模型 (PLM) 具有先进的生物预测,但存在"黑子"性质,限制了透明度和可解释性.
- 解释PLM学习的特征对于人类-AI协作和理解生物机制至关重要.
- 现有的特征解释方法通常需要监督,这可能是劳动密集型的,可能无法捕获所有相关信息.
研究的目的:
- 开发一种无监督的方法,从蛋白质水平和氨基酸水平的PLM表示中提取可解释的特征.
- 评估提取的稀疏特征的生物相关性和解释性.
- 提高PLM在生物应用中的安全性,信任性和可解释性.
主要方法:
- 利用从自然语言处理中调整的稀疏自编码器 (SAEs) 和转码器,从ESM2 PLM中提取特征.
- 在完全无监督的情况下进行特征提取,而不依赖外部的生物标签或探针.
- 利用Anthropic的Claude AI来自动解释提取的稀疏特征.
主要成果:
- SAE成功地从ESM2.2的蛋白质水平和氨基酸水平表示中提取了可解释的稀疏特征.
- 许多提取的稀疏特征显示了与基因本体学 (GO) 术语在所有层次层次上的强烈关联.
- 自动解释确定了与特定蛋白质家族 (例如,NAD激酶,PTH) 相应的特征,功能 (例如,甲基转移酶活性) 和感官感知.
结论:
- SAE提供了一种强大的无监督方法,可以在PLM表示中解开生物相关信息.
- 提取的稀疏特征比原始的PLM神经元更容易解释,有助于发现生物洞察力.
- 这项工作显著提高了PLM的可解释性和可信度,为更深入的生物学理解铺平了道路.
更多相关视频
相关概念视频
Conservation of Protein Domains Over Different Proteins
11.3K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
11.3K
Improving Translational Accuracy
11.9K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.9K
Leaky Scanning
5.2K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.2K
Protein Networks
2.4K
2.4K
Conservation of Protein Domains
3.2K
3.2K
Protein-Protein Interfaces
3.8K
3.8K


