Related Experiment Video
Updated: Sep 11, 2025

Author Spotlight: Streamlining Protein Target Prediction and Validation via Molecular Docking and CETSA
Published on: February 23, 2024
Leveraging large language models for literature-driven prioritization of protein binding pockets
Roman Stratiichuk1,2, Mykola Melnychenko1, Ihor Koleiev1,3
1Receptor.AI Inc., London N1 7GU, United Kingdom.
This study introduces a new method combining geometric pocket detection and large language models (LLMs) to identify protein binding pockets for drug discovery. This approach automates residue extraction, overcoming manual bottlenecks.
Area of Science:
- Computational chemistry
- Bioinformatics
- Drug discovery
Background:
- Identifying protein binding pockets is crucial for small-molecule drug discovery.
- Current manual methods for pocket identification are time-consuming, error-prone, and inefficient.
- Manual curation of binding site data hinders high-throughput drug discovery.
Purpose of the Study:
- To develop a novel, automated approach for identifying and prioritizing protein binding pockets.
- To leverage large language models (LLMs) for extracting residue-level information from scientific literature.
- To improve the efficiency and accuracy of small-molecule drug discovery pipelines.
Main Methods:
- Combined geometric pocket detection (Fpocket) with LLMs for pocket identification and validation.
- Utilized LLMs with fine-tuned prompts to extract residue information from published experimental data.
- Developed and employed a curated benchmark dataset for training and evaluating LLM performance.
Main Results:
- Successfully validated candidate pockets against published experimental data using LLMs.
- Demonstrated LLM's capability in assessing paper relevance and extracting crucial pocket information.
- Established a novel, automated workflow for protein binding pocket identification.
Conclusions:
- The combined geometric and LLM approach offers a significant advancement in automating protein binding pocket identification.
- This method addresses the limitations of manual curation, accelerating drug discovery.
- The developed dataset and methodology are publicly available to facilitate further research.
More Related Videos
06:50Author Spotlight: A Computational Approach to Decipher Amino Acid Preferences in Multispecific Protein-Protein Interactions
Published on: January 26, 2024
08:49Incorporating Target Protein Structure Flexibility and Dynamics in Computational Drug Discovery Using Ensemble-Based Docking Analysis
Published on: June 20, 2025
Related Concept Videos
Ligand Binding Sites
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Protein-protein Interfaces
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Ligand Binding and Linkage
Protein Organization
The primary structure of a protein is its amino acid sequence....
The Equilibrium Binding Constant and Binding Strength