Related Experiment Video
Updated: Jan 6, 2026

Split Point Analysis and Uncertainty Quantification of Thermal-Optical Organic/Elemental Carbon Measurements
Published on: September 7, 2019
A quantitative uncertainty metric controls error in neural network-driven chemical discovery
Jon Paul Janet1, Chenru Duan1,2, Tzuhsiung Yang1
1Department of Chemical Engineering , Massachusetts Institute of Technology , Cambridge , MA 02139 , USA . Email: hjkulik@mit.edu ; Tel: +1-617-253-4584.
Machine learning models accelerate chemical discovery. A new, low-cost uncertainty metric, "distance to available data in latent space," accurately identifies when molecules are outside the model's applicability domain.
Area of Science:
- Computational chemistry
- Materials science
- Artificial intelligence in drug discovery
Background:
- Machine learning (ML) models, including artificial neural networks, offer rapid compound characterization, complementing traditional high-throughput screening.
- Accurate identification of out-of-domain predictions is crucial for reliable large-scale chemical space exploration using ML.
- Existing uncertainty metrics for neural networks are often computationally expensive or require complex feature engineering.
Purpose of the Study:
- To introduce a novel, low-cost uncertainty metric for neural network ML models.
- To enable reliable chemical space exploration by quantifying model applicability domain.
- To provide a method for predictive error control in chemical discovery.
Main Methods:
- Developed a new uncertainty metric based on the distance to available data in the latent space of a neural network.
- Applied the metric to both inorganic and organic chemistry datasets.
- Compared the performance of the new metric against established uncertainty quantification methods.
Main Results:
- The proposed latent distance metric is a low-cost and quantitative measure of uncertainty.
- This approach demonstrates calibrated performance exceeding widely used uncertainty metrics.
- The metric is applicable to models of increasing complexity without additional computational cost.
Conclusions:
- The distance to available data in latent space serves as an effective uncertainty metric for ML models in chemistry.
- This method facilitates predictive error control and aids in identifying valuable data points for active learning strategies.
- The approach enhances the reliability and scalability of ML-driven chemical discovery.
More Related Videos
15:05Functional Evaluation of Biological Neurotoxins in Networked Cultures of Stem Cell-derived Central Nervous System Neurons
Published on: February 5, 2015
08:47Experimental Quantification of Interactions Between Drug Delivery Systems and Cells In Vitro: A Guide for Preclinical Nanomedicine Evaluation
Published on: September 28, 2022
Related Concept Videos
Uncertainty: Overview
Propagation of Uncertainty from Systematic Error
Uncertainty in Measurement: Accuracy and Precision
Propagation of Uncertainty from Random Error
Uncertainty: Confidence Intervals
Data Validation
Key parameters for method validation include: