Related Experiment Video
Updated: Nov 21, 2025

06:19
Constructing and Visualizing Models using Mime-based Machine-learning Framework
Published on: July 22, 2025
1.6K
Dataset's chemical diversity limits the generalizability of machine learning predictions
Marta Glavatskikh1,2, Jules Leguy1, Gilles Hunault1,3
1LERIA, University of Angers, 2 Bd Lavoisier, 49045, Angers, France.
Journal of Cheminformatics
|January 12, 2021
Summary
A new dataset, PC9, offers greater chemical diversity than the standard QM9 dataset for machine learning (ML) models. Models trained on PC9 generalize better to other datasets, improving ML predictions for chemical properties.
Area of Science:
- Computational chemistry
- Materials science
- Chemical informatics
Background:
- The QM9 dataset is a benchmark for machine learning (ML) in predicting chemical properties, derived from GDB's chemical space exploration.
- Recent ML models achieve accuracy comparable to Density Functional Theory (DFT) calculations for molecular predictions.
- Generalizing ML models requires testing on diverse, real-world datasets.
Purpose of the Study:
- Introduce PC9, a new dataset equivalent to QM9 but with enhanced chemical diversity.
- Evaluate the performance of ML models on both QM9 and PC9 datasets.
- Assess the generalization capabilities of models trained on PC9.
Main Methods:
- Utilized Kernel Ridge Regression, Elastic Net, and the SchNet neural network model.
- Performed statistical analysis of bonding distances and chemical functions within PC9.
- Trained and tested ML models on both the QM9 and PC9 datasets.
Main Results:
- The PC9 dataset exhibits greater chemical diversity compared to QM9.
- While QM9 subset yielded higher overall energy prediction accuracy, PC9-trained models demonstrated superior generalization.
- Models trained on PC9 showed a stronger ability to predict energies for the other dataset.
Conclusions:
- PC9 represents a valuable, more diverse alternative to QM9 for ML model development.
- ML models trained on PC9 exhibit improved generalization, crucial for real-world chemical property prediction.
- The findings support the use of PC9 for robust testing and validation of ML models in chemistry.
Related Concept Videos
Survival Tree
245
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
245
Generalization, Discrimination, and Extinction
1.1K
Generalization, discrimination, and extinction are key concepts in operant conditioning that influence how behaviors are learned and maintained.
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
Generalization occurs when a behavior reinforced in one context is performed in similar situations. For instance, a student who studies diligently for calculus and receives excellent grades might apply the same study habits to psychology and history, expecting similar results. Generalization shows how learning in one setting can influence behavior in...
1.1K
Mechanistic Models: Compartment Models in Individual and Population Analysis
153
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
153
Prediction Intervals
2.7K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.7K
Genetic Drift
42.1K
Natural selection—probably the most well-known evolutionary mechanism—increases the prevalence of traits that enhance survival and reproduction. However, evolution does not merely propagate favorable traits, nor does it always benefit populations.
42.1K
Variability: Analysis
285
Measures of variability are statistical metrics that reveal the dispersion pattern within a dataset. They are pivotal in biostatistics, providing insights into the heterogeneity within health and biological data. Variability signifies the degree to which data points diverge from one another, helping researchers understand the potential range of values and associated uncertainty within the data.
The range is a simple measure of variability, indicating the difference between the highest and...
The range is a simple measure of variability, indicating the difference between the highest and...
285
