Uncertainty Quantification and Flagging of Unreliable Predictions in Predicting Mass Spectrometry-Related Properties
Dmitriy D Matyushin1, Ivan A Burov1, Anastasia Yu Sholokhova1
1A.N. Frumkin Institute of Physical Chemistry and Electrochemistry, Russian Academy of Sciences, 31 Leninsky Prospect, GSP-1, 119071 Moscow, Russia.
This study introduces a novel method to estimate prediction uncertainty in molecular identification. By combining ensemble model spread, molecular similarity, and data clustering, researchers can better assess the reliability of predictions in metabolomics and chromatography.
Area of Science:
- Analytical Chemistry
- Computational Chemistry
- Cheminformatics
Background:
- Mass spectral identification is crucial in metabolomics, often enhanced by predicting molecular properties like chromatographic retention.
- Current machine learning models for prediction typically lack robust uncertainty estimation.
- Accurate uncertainty quantification is vital for reliable molecular identification and data interpretation.
Purpose of the Study:
- To develop and validate a method for assessing prediction uncertainty in molecular property prediction.
- To improve the reliability of machine learning models used in mass spectral identification and metabolomics.
- To integrate multiple uncertainty indicators into a unified assessment framework.
Main Methods:
- Utilized ensemble model prediction spread as an uncertainty indicator.
- Employed Euclidean distance of molecular descriptors to quantify molecular similarity.
- Incorporated data set clustering to identify potential outliers or uncertain data points.
- Developed classification models using these factors to predict prediction reliability.
Main Results:
- Achieved area under the receiver operating curve (AUC) values between 0.73 and 0.82 for predicting unreliable predictions.
- Demonstrated the effectiveness of combining ensemble spread, molecular similarity, and clustering for uncertainty assessment.
- Successfully predicted whether a prediction falls within the worst 15% of outcomes across various chromatographic and ion mobility tasks.
Conclusions:
- The proposed method effectively quantifies prediction uncertainty in molecular property prediction tasks.
- Integrating ensemble spread, molecular similarity, and clustering provides a comprehensive approach to uncertainty estimation.
- This work enhances the trustworthiness of machine learning predictions in analytical chemistry and metabolomics.
Related Concept Videos
Mass Spectrometry: Complex Analysis
GC–MS is a powerful hyphenated method commonly used in forensics and environmental...
Mass Spectrometry: Overview
High-Resolution Mass Spectrometry (HRMS)
Peptide Identification Using Tandem Mass Spectrometry
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
Mass Spectrum: Interpretation
To...
Mass Analyzers: Overview


