Related Experiment Video
Updated: Apr 3, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Yield Prediction of Organic Reactions in Biased Data Sets via Positive-Unlabeled Learning
Florian Boser1, Jan C Spies1, Frank Glorius1
1Organisch-Chemisches-Institut, Universität Münster, Corrensstraße 36, 48149 Münster, Germany.
Abstract:
The vast reaction data within scientific literature represents a rich resource for training predictive machine learning models. However, this resource is fundamentally compromised by a pervasive selection and reporting bias, resulting in imbalanced data sets. In this work, we introduce "Positivity is All You Need" (PAYN), a machine learning framework that addresses this data-scarcity problem by learning directly from biased, positive-only data. PAYN leverages a spy-based positive-unlabeled (PU) learning strategy, treating reported high-yielding reactions as the "positive" class and the vast, unexplored chemical space as the "unlabeled" class. To validate our approach, we simulated literature bias on fully labeled high-throughput experimentation (HTE) data sets, including Ni-catalyzed borylations, Buchwald-Hartwig and Suzuki-Miyaura couplings. We demonstrated that PAYN significantly improves the performance of models trained on biased data by balancing the data with augmented negative data points. This work establishes a robust strategy for leveraging biased data, paving a path toward more scalable and accessible data-driven strategies for accelerating synthesis design, optimization, and chemical discovery.
Related Concept Videos
Predicting Reaction Outcomes
Predicting Products: SN1 vs. SN2
With increased substitution on the alkyl halide,...
Predicting Products: Substitution vs. Elimination
The following factors can influence the mechanisms competing against each other:
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Classification of Titrimetric Analysis Based on Reaction Types
Titrations between an acid and a base lead to neutralization reactions that form...
Polymer Classification: Stereospecificity