Related Experiment Video
Updated: Sep 9, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Empowering Reactivity Predictions through Noise-Based Data Augmentation
Julian A Hueffel1, Quentin P Bindschaedler1, Francesco Sala1
1Institute of Organic Chemistry, RWTH Aachen University, Landoltweg 1, 52074 Aachen, Germany.
None:
Data scarcity is a key obstacle in the pursuit to capitalize on the predictive powers of the A.I. in problems related to bond-making and -breaking at the molecular level. While the generation of artificial data from real data points (known as "data augmentation") is a widely pursued strategy employed in, for example, image or speech recognition or health data evaluations, among others, to artificially expand existing data sets, it is currently unknown whether this strategy is applicable to reactivity problems at the molecular level, where predictive models are exquisitely sensitive to steric, electronic, and structural nuances. Here, we systematically evaluated the power of data augmentation for a diverse set of reactivity questions ranging from the prediction of activation barriers to stereoselectivities of catalytic transformations. We demonstrate that introducing Gaussian noise to existing data points, which is completed in under a second for a full data set, can dramatically enhance the predictive performance. It can enable model training in low-data regimes where otherwise no meaningful model could be built and achieves accuracy comparable to models built on full data sets, while needing only a fraction of the data. The approach substantially lowers the number of necessary experiments (by 20-50%), conserving time, energy, and resources, while advancing the integration of machine learning in molecular reactivity challenges.
Related Concept Videos
Predicting Reaction Outcomes
Amplifying Signals via Enzymatic Cascade
Survival Tree
Building a Survival Tree
Constructing a...
Standard Entropy Change for a Reaction
Randomized Experiments
Simple randomization
Simple...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
