Related Experiment Video
Updated: Sep 18, 2025

Optimization of Synthetic Proteins: Identification of Interpositional Dependencies Indicating Structurally and/or Functionally Linked Residues
Published on: July 14, 2015
ProtTeX: Structure-In-Context Reasoning and Editing of Proteins with Large Language Models
Zicheng Ma1,2, Chuanliu Fan3, Zhicong Wang3
1Changping Laboratory, Beijing 102200, China.
ProtTeX integrates protein sequences, structures, and text into a unified token space. This enables large language models (LLMs) to perform structure-aware protein analysis and generation, significantly improving prediction accuracy.
Area of Science:
- Computational biology
- Artificial intelligence in protein science
Background:
- Large language models (LLMs) excel in molecular science, primarily using sequence-based tokenization for small molecules.
- Protein science challenges are often structure-dependent, yet current LLMs lack structure-aware tokenization, limiting their capabilities.
- Existing LLMs struggle with comprehensive biomolecular understanding and multimodal generation due to the absence of structural information.
Purpose of the Study:
- Introduce ProtTeX, a novel framework for tokenizing protein sequences, structures, and text into a unified discrete space.
- Enable joint LLM training via Next-Token Prediction for multimodal protein reasoning and generation.
- Empower general LLMs to process protein structures, use them for reasoning, and generate/manipulate structures via text.
Main Methods:
- Developed ProtTeX to tokenize protein sequences, structures, and associated textual information.
- Implemented a unified discrete space for multimodal protein data representation.
- Utilized the Next-Token Prediction paradigm for joint LLM training, enabling structure-aware processing.
Main Results:
- ProtTeX significantly enhances protein function prediction accuracy, achieving a 2-fold improvement over state-of-the-art domain expert models.
- Demonstrated high-quality protein conformational generation and customizable protein design capabilities.
- Showcased the ability of decoder-only LLMs, using standard pipelines, to address diverse protein-related tasks with ProtTeX.
Conclusions:
- ProtTeX bridges the gap between LLMs and structure-dependent protein science challenges.
- The framework facilitates advanced multimodal protein reasoning and generation using standard LLM architectures.
- ProtTeX represents a significant advancement in applying LLMs to complex protein design and analysis problems.
Related Concept Videos
Protein Organization
The primary structure of a protein is its amino acid sequence....
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
Protein and Protein Structures
Protein Complex Assembly
Protein and Protein Structure
A protein's shape is critical to its function. For example, an enzyme...
Protein Complexes with Interchangeable Parts
The SCF ubiquitin ligase is a protein complex of five individual proteins. This complex attaches ubiquitin to other target proteins to mark them for degradation. In order...

