Related Experiment Video
Updated: Dec 22, 2025

Computation of Atmospheric Concentrations of Molecular Clusters from ab initio Thermochemistry
Published on: April 8, 2020
Probabilistic performance estimators for computational chemistry methods: Systematic improvement probability and
Pascal Pernot1, Andreas Savin2
1Institut de Chimie Physique, UMR8000, CNRS, Université Paris-Saclay, 91405 Orsay, France.
Evaluating computational chemistry theories requires new methods beyond standard error rankings. This study introduces a systematic improvement probability score and robust statistical indicators to assess ranking reliability, addressing benchmark dataset limitations.
Area of Science:
- Computational chemistry
- Theoretical chemistry
- Chemical physics
Background:
- Standard mean unsigned error rankings in computational chemistry are limited by non-normal error distributions and trends.
- Existing complementary statistics like error quantiles and prediction uncertainty partially address these limitations.
- Uncertainty arising from incomplete benchmark datasets is often overlooked in method evaluation.
Purpose of the Study:
- To introduce a novel scoring system, the systematic improvement probability, for direct system-wise comparison of absolute errors.
- To develop robust statistical indicators, inversion probability (P_inv) and ranking probability matrix (P_r), to quantify ranking uncertainty.
- To highlight the importance of correlations between error sets in statistical comparisons.
Main Methods:
- Direct system-wise comparison of absolute errors to calculate systematic improvement probability.
- Development of robust statistical indicators: inversion probability (P_inv) and ranking probability matrix (P_r).
- Analysis of correlations between error sets to assess their impact on ranking statistics.
Main Results:
- The proposed systematic improvement probability offers a new metric for evaluating computational chemistry methods.
- P_inv and P_r provide essential measures of ranking robustness against benchmark dataset incompleteness.
- Correlations between error sets significantly influence the reliability of statistical comparisons.
Conclusions:
- The standard mean unsigned error is insufficient for robustly ranking computational chemistry methods.
- Novel statistical approaches, including systematic improvement probability and robust indicators (P_inv, P_r), enhance the evaluation of theoretical models.
- Accurate assessment of computational chemistry theories necessitates considering error distributions, dataset incompleteness, and error correlations.
Related Concept Videos
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Estimation of the Physical Quantities
Uncertainty: Overview
Propagation of Uncertainty from Systematic Error
Propagation of Uncertainty from Random Error
Mechanistic Models: Compartment Models in Individual and Population Analysis

