Related Experiment Video
Updated: Aug 11, 2025

03:37
Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
Published on: March 1, 2024
832
Blinded Predictions and Post Hoc Analysis of the Second Solubility Challenge Data: Exploring Training Data and
Jonathan G M Conn1, James W Carter1, Justin J A Conn1
1Department of Pure and Applied Chemistry, University of Strathclyde, Thomas Graham Building, 295 Cathedral Street, Glasgow G1 1XL, U.K.
Journal of Chemical Information and Modeling
|February 9, 2023
Summary
Predicting chemical solubility from molecular structure is crucial. Machine learning models, especially graph convolutional neural networks, show promise, but high-quality, relevant training data is key for accurate solubility predictions.
Area of Science:
- Chemical Sciences
- Computational Chemistry
- Drug Discovery
Background:
- Accurate prediction of solubility from molecular structure is a significant challenge in chemical sciences.
- The American Chemical Society's "Second Solubility Challenge" (2019) aimed to assess the state-of-the-art in solubility prediction.
- Competitors submitted blinded predictions for 132 drug-like molecules.
Purpose of the Study:
- To develop and report two novel machine learning models for solubility prediction submitted to the 2019 challenge.
- To evaluate the impact of advanced algorithms and larger datasets on prediction accuracy by comparing original models with deep learning models.
- To analyze the performance differences between models based on feature sets and training data.
Main Methods:
- Development of two traditional machine learning models using inexpensive molecular descriptors and a small training set (300 molecules).
- Training and evaluation of deep learning models, including graph convolutional neural networks, on larger datasets (2999 and 5697 molecules).
- Comparison of prediction accuracy (RMSE) between original and deep learning models on the challenge dataset.
Main Results:
- Several algorithms achieved near state-of-the-art performance.
- A graph convolutional neural network model yielded the best performance with an RMSE of 0.86 log units.
- Systematic differences in model performance were observed based on feature sets and training data quality and relevance.
Conclusions:
- Machine learning models can achieve high accuracy in predicting solubility.
- The quality and relevance of training data are critical for accurate solubility prediction.
- Challenges remain in modeling complex chemical spaces with sparse training data, highlighting the need for careful data selection and methodological improvements.
Related Concept Videos
Survival Tree
130
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
130
Predicting Reaction Outcomes
8.5K
Kinetics describes the rate and path by which a reaction occurs. In contrast, thermodynamics deals with state functions and describes the properties, behavior, and components of a system. It is not concerned with the path taken by the process and cannot address the rate at which a reaction occurs. Although it does provide information about what can happen during a reaction process, it does not describe the detailed steps of what appears on an atomic or a molecular level. On the other hand,...
8.5K

