Related Experiment Video
Updated: Jan 16, 2026

Computation of Atmospheric Concentrations of Molecular Clusters from ab initio Thermochemistry
Published on: April 8, 2020
Fast and Interpretable Machine Learning Modeling of Atmospheric Molecular Clusters
Lauri Seppäläinen1, Jakub Kubečka2, Jonas Elm2
1Department of Computer Science, University of Helsinki, Pietari Kalmin katu 5, 00560 Helsinki, Finland.
None:
Understanding how atmospheric molecular clusters form and grow is key to resolving one of the biggest uncertainties in climate modeling: the formation of new aerosol particles. While quantum chemistry offers accurate insights into these early-stage clusters, its steep computational costs limit large-scale exploration. In this work, we present a fast, interpretable, and surprisingly powerful alternative: the k-nearest neighbor (k-NN) regression model. By leveraging chemically informed distance metrics, including a kernel-induced metric and one learned via metric learning for kernel regression (MLKR), we show that simple k-NN models can rival more complex kernel ridge regression (KRR) models in accuracy while reducing computational time by orders of magnitude. We perform this comparison with the well-established Faber-Christensen-Huang-Lilienfeld (FCHL19) molecular descriptor; however, other descriptors (e.g., FCHL18, MBDF, and CM) can be shown to have similar performance. Applied to both simple organic molecules in the QM9 benchmark set and large data sets of atmospheric molecular clusters (sulfuric acid-water and sulfuric-multibase-base systems), our k-NN models achieve near-chemical accuracy, scale seamlessly to data sets with over 250,000 entries, and even appears to extrapolate to larger unseen clusters with minimal error (often nearing 1 kcal/mol). With built-in interpretability and straightforward uncertainty estimation, this work positions k-NN as a potent tool for accelerating discovery in atmospheric chemistry and beyond.
Related Concept Videos
Predicting Molecular Geometry
Molecular Models
Molecular Comparison of Gases, Liquids, and Solids
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Distribution of Molecular Speeds
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...

