Related Experiment Videos
Yield Smarter, Not Harder: Good Practices for Machine Learning of Reaction Outcomes
Idil Ismail1, Gregory A Landrum1, Sereina Riniker1
1Department of Chemistry and Applied Biosciences, ETH Zurich, Vladimir-Prelog-Weg 2, Zurich8093, Switzerland.
Abstract:
Reaction yield prediction is a longstanding challenge in synthetic chemistry, with broad implications for route planning, scalability, and high-throughput experimentation (HTE). While recent machine learning (ML) approaches have demonstrated promise in modeling reactivity, they often use complex descriptors or deep architectures that are computationally expensive and limit interpretability and scalability. Here, we assess how much information is stored in simpler descriptors and whether model accuracy is improved by increasing the complexity of the descriptors. Using classical ML models trained on descriptors with different complexity levels, we benchmark predictive performance on four publicly available HTE data sets covering three diverse reaction data sets: Buchwald-Hartwig (BH) amination, Suzuki-Miyaura (SM) coupling, and the silicon-amine protocol (SLAP). Our evaluation furthermore discusses (1) generalization via component-wise data splitting, (2) robustness through external validation across data sets, and (3) performance across asymmetric yield distributions characteristic of HTE data. Contrary to conventional expectations, we find that simpler models with interpretable features can achieve competitive performance under rigorous validation protocols. Based on our findings, we formulate good practices for future studies in this area. For example, comparison to low-cost baseline models should become a requirement for future ML studies for reaction-yield prediction.
Related Concept Videos
Predicting Reaction Outcomes
Methods of Medium Optimization
Law of Effect
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle boxes...
Standard Entropy Change for a Reaction
Response Surface Methodology
The process of RSM involves several key steps:
Decision Making: P-value Method
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can have a...