Related Experiment Video
Updated: May 13, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Exploring BERT for Reaction Yield Prediction: Evaluating the Impact of Tokenization, Molecular Representation, and
Adrian Krzyzanowski1, Stephen D Pickett1, Peter Pogány1
1GSK Medicines Research Centre, Gunnels Wood Road, Stevenage SG1 2NY, U.K.
Machine learning models can predict chemical reaction yields effectively. Key factors include tokenization and pretraining data, with adversarial training enhancing model robustness for synthetic chemistry applications.
Area of Science:
- Computational chemistry
- Machine learning in chemistry
- Synthetic chemistry
Background:
- Predicting reaction yields is crucial but challenging in synthetic chemistry.
- High-throughput experimentation (HTE) data provides valuable information for model training.
- BERT-based models show promise for reaction yield prediction.
Purpose of the Study:
- To systematically evaluate factors influencing BERT-based yield prediction models.
- To assess the impact of tokenization, molecular representation, and pretraining data.
- To investigate the effectiveness of adversarial training for improved model performance.
Main Methods:
- Utilized BERT-based models for yield prediction of Buchwald-Hartwig and Suzuki-Miyaura coupling reactions.
- Evaluated various tokenization methods (BPE, SentencePiece, WordPiece) and molecular representations (SMILES, DeepSMILES, SELFIES, etc.).
- Compared performance using different pretraining data sizes and explored artificially generated datasets, alongside adversarial training.
Main Results:
- Molecular representation had minimal impact; BPE and SentencePiece tokenization performed best.
- Smaller pretraining datasets (<100 K reactions) yielded comparable results to larger ones.
- Hybrid pretraining sets (real + artificial data) and adversarial training improved model robustness and generalizability.
Conclusions:
- Tokenization and pretraining strategies significantly impact ML model performance for chemical yield prediction.
- Artificially generated data can serve as an effective surrogate for real-world datasets.
- Adversarial training offers a novel approach to enhance the robustness and applicability of ML models in synthetic chemistry.
More Related Videos
06:19Integration of Animal Behavioral Assessment and Convolutional Neural Network to Study Wasabi-Alcohol Taste-Smell Interaction
Published on: August 16, 2024
09:47Author Spotlight: Advancing Alzheimer's Research – Exploring Early Detection and Multi-Omics Approaches
Published on: December 15, 2023
Related Concept Videos
Predicting Reaction Outcomes
Reaction Quotient
Reaction Yield
Rate-Determining Steps
In a multistep reaction mechanism, one of the elementary steps progresses significantly slower than the others. This slowest step is called the rate-limiting step (or rate-determining step). A reaction cannot proceed faster than its slowest step, and hence, the rate-determining step limits the overall reaction rate.
The concept of rate-determining step can be understood from the analogy of a 4-lane freeway with a short-stretch of traffic-bottleneck caused due to...
Classification of Titrimetric Analysis Based on Reaction Types
Titrations between an acid and a base lead to neutralization reactions that form...
Multi-Step Reactions