Related Experiment Video
Updated: Jun 7, 2025

Computation of Atmospheric Concentrations of Molecular Clusters from ab initio Thermochemistry
Published on: April 8, 2020
k-Means Clustering in Fingerprint-Based Configuration Selection for Fitting Interatomic Potentials.
Miroslav Lebeda1,2, Jan Drahokoupil1,3, Ludvík Löbel3
1Department of Physics, Faculty of Mechanical Engineering, Czech Technical University in Prague, Technická 4, Prague 6 16607, Czech Republic.
K-means clustering efficiently selects distinct atomistic configurations, improving interatomic potential accuracy for materials science simulations. This method requires fewer configurations than random selection for accurate energy and force predictions.
Area of Science:
- Computational Materials Science
- Atomistic Simulations
- Machine Learning in Chemistry
Background:
- Accurate interatomic potentials are crucial for molecular dynamics simulations.
- Fitting potentials to quantum mechanical data (like DFT) requires a representative set of atomic configurations.
- Traditional methods often rely on random sampling, which can be inefficient.
Purpose of the Study:
- To develop a more efficient method for selecting representative atomic configurations for fitting interatomic potentials.
- To improve the accuracy and reduce the number of configurations needed for fitting potentials.
- To compare the efficacy of k-means clustering against random selection.
Main Methods:
- Applying k-means clustering to atomistic configuration fingerprints.
- Utilizing CrystalNN model and radial distribution function (RDF) for fingerprinting.
- Fitting embedded-atom method (EAM) potentials for titanium using selected configurations.
- Employing t-distributed stochastic neighbor embedding (t-SNE) for dimensionality reduction and visualization.
Main Results:
- K-means clustering significantly improves the accuracy of fitting interatomic potentials (energies and forces) compared to random selection.
- Fewer configurations (around 30) selected by k-means are sufficient to describe a larger dataset (1800 configurations).
- K-means clustering effectively identifies and utilizes configurations with similar atomic environments, even those with vacancies, which random selection misses.
Conclusions:
- K-means clustering offers a superior strategy for selecting configurations in materials modeling.
- This approach leads to more precise and reliable interatomic potentials with reduced computational cost.
- The findings suggest potential information redundancy in atomistic configurations, which k-means can exploit.
More Related Videos
08:59Determination of Aggregate Surface Morphology at the Interfacial Transition Zone ITZ
Published on: December 16, 2019
09:17Structure-Based Simulation and Sampling of Transcription Factor Protein Movements along DNA from Atomic-Scale Stepping to Coarse-Grained Diffusion
Published on: March 1, 2022
Related Concept Videos
IR Frequency Region: Fingerprint Region
¹H NMR: Interpreting Distorted and Overlapping Signals
As Δν decreases and the signals move closer, the doublets appear increasingly distorted. The intensities of the inner lines increase at the cost of those of the outer lines as the signals are...
Hückel's Rule Diagram of π MOs: Frost Circle
A Frost circle is constructed by drawing a polygon whose number of edges is equal to the number of carbons of the given cyclic system, with one of the vertices pointing down. Then, a circle is drawn enclosing the polygon so...