Related Experiment Video
Updated: May 10, 2025

Author Spotlight: Streamlining Protein Target Prediction and Validation via Molecular Docking and CETSA
Published on: February 23, 2024
Molecular property prediction using pretrained-BERT and Bayesian active learning: a data-efficient approach to drug
Muhammad Arslan Masood1, Samuel Kaski2,3, Tianyu Cui2
1Department of Computer Science, Aalto University, Espoo, Finland. arslan.masood@aalto.fi.
This study enhances drug discovery by using pretrained BERT models with active learning to improve molecule selection. This approach identifies toxic compounds more efficiently, requiring fewer experimental tests.
Area of Science:
- Computational chemistry
- Machine learning in drug discovery
- Bioinformatics
Background:
- Prioritizing compounds for testing is crucial in drug discovery.
- Active learning typically uses only labeled data, ignoring unlabeled molecular data.
- This limits predictive performance and molecule selection efficiency.
Purpose of the Study:
- To improve active learning in drug discovery by integrating pretrained transformer models.
- To address the limitation of fully supervised active learning by leveraging unlabeled molecular data.
- To enhance molecule selection and predictive performance in compound prioritization.
Main Methods:
- Integrated a transformer-based BERT model, pretrained on 1.26 million compounds, into an active learning pipeline.
- Separated representation learning from uncertainty estimation using pretrained molecular representations.
- Conducted experiments on Tox21 and ClinTox datasets.
Main Results:
- Achieved equivalent toxic compound identification with 50% fewer iterations compared to conventional active learning.
- Demonstrated that pretrained BERT representations create a structured embedding space for reliable uncertainty estimation with limited labeled data.
- Showcased improved model performance and acquisition efficiency in drug discovery.
Conclusions:
- High-quality molecular representations are fundamental to active learning success in drug discovery.
- The proposed framework integrates pretrained transformer models with Bayesian active learning for efficient screening.
- This approach provides a scalable foundation for optimizing compound prioritization in pharmaceutical research.
More Related Videos
06:50Author Spotlight: A Computational Approach to Decipher Amino Acid Preferences in Multispecific Protein-Protein Interactions
Published on: January 26, 2024
08:31Biosensor-based High Throughput Biopanning and Bioinformatics Analysis Strategy for the Global Validation of Drug-protein Interactions
Published on: December 1, 2020
Related Concept Videos
Structure-Activity Relationships and Drug Design
SAR studies the intricate relationship between a drug's chemical structure and biological activity. It focuses on understanding how modifications to a drug's structure can influence...
Drug Discovery: Overview
Ligand Binding Sites
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Prodrugs
Prodrugs help overcome...
Principles of Drug Action
Drugs can be agonists or antagonists. Like the endogenous ligands, agonists always bind and activate the target to produce a cellular response. Agonist binding induces a conformational change which in turn...
The Two-State Receptor Model
The binding affinity of a drug determines its interaction with...