Related Experiment Video
Updated: Jun 12, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Aligning large language models and geometric deep models for protein representation
Dong Shu1, Bingbing Duan2, Kai Guo3
1Northwestern University, Computer Science Department, Evanston, IL 60201, USA.
Abstract:
In this study, we explore the alignment of multimodal representations between large language models (LLMs) and geometric deep models (GDMs) in the protein domain. We comprehensively evaluate three LLMs with four protein-specialized GDMs. Our work examines alignment factors from both model and protein perspectives, identifying challenges in current alignment methodologies and proposing strategies to improve the alignment process. Experimental results reveal that GDMs incorporating both graph and 3D structural information align better with LLMs, larger LLMs demonstrate improved alignment capabilities, and protein rarity significantly impacts alignment performance. We also find that increasing GDM embedding dimensions, using two-layer projection heads, and fine-tuning LLMs on protein-specific data substantially enhance alignment quality. Last, we demonstrate that improved alignment correlates with better downstream performance and reduced hallucination in protein-focused multimodal LLMs.
More Related Videos
Related Concept Videos
Protein Organization
The primary structure of a protein is its amino acid sequence....
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Protein-protein Interfaces

