Related Experiment Video
Updated: Jun 1, 2025

16:41
A Protocol for Computer-Based Protein Structure and Function Prediction
Published on: November 3, 2011
68.5K
A novel computational machine learning pipeline to quantify similarities in 3D protein structures.
Shreyas U Hirway1, Xiao Xu2, Fan Fan1
1Janssen Research & Development, LLC, La Jolla, CA 92121, United States.
Summary
Selecting the best animal model for drug development is crucial. This study introduces a machine learning pipeline that uses 3D protein structures to identify optimal animal models, improving preclinical research accuracy.
Area of Science:
- Biomedical research
- Computational biology
- Pharmacology
Background:
- Animal models are essential in drug development for preclinical testing.
- Selecting appropriate animal models requires assessing target biology and protein similarity to humans.
- Current methods rely on sequence comparison, neglecting crucial 3D structural information.
Purpose of the Study:
- To develop a novel machine learning pipeline for selecting optimal animal models in drug development.
- To enhance the accuracy of cross-species protein similarity assessment by incorporating 3D structural data.
- To provide a quantitative method for identifying the most relevant animal species for specific drug targets.
Main Methods:
- Utilized a machine learning pipeline incorporating 3D structure-based features from the Protein Data Bank.
- Integrated nominal data from UNIPROT and bioactivity data from ChEMBL, matched for human and animal data.
- Employed an XGBoost regression model to calculate target similarity scores and predict optimal animal species, including validation with AlphaFold-derived targets.
Main Results:
- Developed a quantitative pipeline to calculate cross-species protein similarity scores based on 3D structural features.
- Successfully identified optimal animal species for drug targets by comparing similarity scores.
- Demonstrated the pipeline's applicability using alternative protein structure data (AlphaFold) and grouping targets by phenotype.
Conclusions:
- The novel 3D structure-based machine learning pipeline significantly improves the selection of animal models for drug development.
- This approach offers a more biologically relevant and accurate method for assessing cross-species protein similarity compared to traditional sequence-based methods.
- The pipeline facilitates the identification of optimal animal models, thereby enhancing the efficiency and reliability of preclinical drug testing and development.

