大气分子团的快速和可解释的机器学习建模
Lauri Seppäläinen1, Jakub Kubečka2, Jonas Elm2
1Department of Computer Science, University of Helsinki, Pietari Kalmin katu 5, 00560 Helsinki, Finland.
The journal of physical chemistry. A
|January 15, 2026
概括
我们开发了一种快速的k-最近邻居 (k-NN) 模型,用于预测大气分子团的特性. 与量子化学相比,这种方法显著降低了计算成本,有助于气候建模和气溶形成研究.
科学领域:
- 大气化学 大气化学
- 计算化学计算化学
- 气候科学 气候科学
背景情况:
- 了解大气分子团的形成对于准确的气候建模和预测新的气溶颗粒的形成至关重要.
- 当前的量子化学方法提供了高精度,但在计算上昂贵,限制了大规模研究.
- 开发高效的计算模型对于推进大气化学研究至关重要.
研究的目的:
- 为研究大气分子团的计算密集型方法提供快速,可解释和准确的替代方案.
- 通过使用化学信息的距离指标来评估k-最近邻居 (k-NN) 回归模型的性能.
- 为了证明k-NN模型对大气系统的可扩展性和预测能力.
主要方法:
- 采用了k-最近邻居 (k-NN) 回归模型.
- 利用化学信息的距离指标,包括内核诱导的和用于内核回归 (MLKR) 的指标的指标学习.
- 使用FCHL19分子描述器和其他描述器对核回归 (KRR) 的k-NN性能进行比较.
- 将模型应用于QM9基准数据和硫酸-水和硫酸-多基基集群的大数据集.
主要成果:
- k-NN模型的准确性与KRR模型相比,但计算时间缩短了数量级.
- 模型在基准和大型大气集群数据集 (>250,000条目) 上都显示出近化学准确性.
- k-NN方法在推断到更大,未见的集群时显示出最小的误差,通常接近1kcal/mol.
结论:
- k-最近邻居 (k-NN) 回归为研究大气分子团提供了一个计算效率高,准确的方法.
- 开发的k-NN模型具有内置的可解释性和不确定性估计,可以加速大气化学的发现.
- 这项工作将k-NN定位为改进气候模型和了解气溶形成过程的强大工具.
相关概念视频
Predicting Molecular Geometry
45.0K
VSEPR Theory for Determination of Electron Pair Geometries
45.0K
Molecular Models
43.5K
Physical models representing molecular architectures of chemical compounds play essential roles in understanding chemistry. The use of molecular models makes it easier to visualize the structures and shapes of atoms and molecules.
43.5K
Molecular Comparison of Gases, Liquids, and Solids
53.6K
Particles in a solid are tightly packed together (fixed shape) and often arranged in a regular pattern; in a liquid, they are close together with no regular arrangement (no fixed shape); in a gas, they are far apart with no regular arrangement (no fixed shape). Particles in a solid vibrate about fixed positions (cannot flow) and do not generally move in relation to one another; in a liquid, they move past each other (can flow) but remain in essentially constant contact; in a gas, they move...
53.6K
Cluster Sampling Method
14.0K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
14.0K
Distribution of Molecular Speeds
5.3K
The motion of molecules in a gas is random in magnitude and direction for individual molecules, but a gas of many molecules has a predictable distribution of molecular speeds. This predictable distribution of molecular speeds is known as the Maxwell-Boltzmann distribution. The distribution of molecular speeds in liquids is comparable to that of gases but not identical and can help to understand the phenomenon of the boiling and vapor pressure of a liquid. Consider that a molecule requires a...
5.3K
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
242
Pharmacokinetic models are mathematical constructs that represent and predict the time course of drug concentrations in the body, providing meaningful pharmacokinetic parameters. These models are categorized into compartment, physiological, and distributed parameter models.
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
242


