Related Experiment Video
Updated: Aug 5, 2025

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
Published on: March 1, 2024
AiZynthTrain: Robust, Reproducible, and Extensible Pipelines for Training Synthesis Prediction Models
Samuel Genheden1, Per-Ola Norrby2, Ola Engkvist1
1Molecular AI, Discovery Sciences, R&D, AstraZeneca Gothenburg, SE-431 83 Mölndal, Sweden.
The AiZynthTrain Python package offers reproducible training for chemical synthesis models, including retrosynthesis and RingBreaker models. New heuristics improve ring system disconnection, enhancing model performance.
Area of Science:
- Computational chemistry
- Machine learning in chemistry
- Chemical synthesis prediction
Background:
- Retrosynthesis is crucial for planning chemical synthesis.
- Existing methods for training synthesis models often lack reproducibility and extensibility.
- Developing robust and adaptable tools for computational chemistry is essential.
Purpose of the Study:
- Introduce the AiZynthTrain Python package for training chemical synthesis models.
- Provide reproducible pipelines for template-based retrosynthesis and RingBreaker models.
- Demonstrate the package's robustness and potential for future extensions.
Main Methods:
- Developed two training pipelines within the AiZynthTrain package.
- Utilized a publicly available reaction dataset from the U.S. Patent and Trademark Office (USPTO).
- Implemented new heuristics to enhance the RingBreaker model's performance on ring system disconnection.
Main Results:
- Created the first end-to-end reproducible retrosynthesis models.
- The RingBreaker model showed significantly improved performance in disconnecting ring systems.
- Demonstrated pipeline robustness through training on a diverse proprietary dataset.
Conclusions:
- AiZynthTrain provides a robust, reproducible, and extensible framework for training synthesis models.
- The implemented heuristics enhance the capability of retrosynthesis prediction, particularly for complex ring systems.
- The framework is poised for future integration of additional synthesis models.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Related Concept Videos
Improving Translational Accuracy
Predicting Reaction Outcomes
Predicting Products: SN1 vs. SN2
With increased substitution on the alkyl halide,...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Predicting Products: Substitution vs. Elimination
The following factors can influence the mechanisms competing against each other: