DONKEY: A Flexible and Accurate Algorithm for Clustering
Jakub Kára1, Kyle Acheson2, Adam Kirrander1
1Physical and Theoretical Chemistry Laboratory, Department of Chemistry, University of Oxford, South Parks Road, Oxford OX1 3QZ, United Kingdom.
Journal of Chemical Theory and Computation
|May 2, 2025
Summary
We developed a new clustering algorithm for analyzing complex photoexcited dynamics simulations. This method accurately identifies distinct reaction pathways without parameter tuning, improving data analysis in computational chemistry.
Area of Science:
- Computational chemistry
- Chemical physics
- Data science
Background:
- Analyzing complex molecular dynamics simulations, especially photoexcited states, presents significant data challenges.
- Existing clustering algorithms often require parameter tuning, introducing bias and limiting applicability to varied datasets.
Purpose of the Study:
- To introduce a novel, parameter-free clustering algorithm for analyzing multidimensional temporal data from nonadiabatic trajectory-based simulations.
- To enhance the identification of distinct reaction pathways in photoexcited molecular systems.
Main Methods:
- Variable kernel density estimation to approximate probability density functions.
- Assignment of data points to local maxima representing cluster centers.
- Merging of clusters to overcome minor density fluctuations, ensuring robust separation.
Main Results:
- The algorithm demonstrates superior performance compared to conventional methods on synthetic datasets.
- Successfully applied to the photoexcited dynamics of the norbornadiene ⇌ quadricyclane molecular photoswitch.
- Identified distinct reaction pathways, showcasing its practical utility.
Conclusions:
- The proposed clustering algorithm offers an accurate and flexible tool for analyzing complex simulation data.
- Its parameter-free nature reduces bias and enhances applicability across diverse scientific domains.
- Facilitates deeper insights into reaction mechanisms in photochemistry and molecular dynamics.
Related Concept Videos
Cluster Sampling Method
11.5K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.5K
Sampling Plans
155
Sampling is a crucial step in analytical chemistry, allowing researchers to collect representative data from a large population. Common sampling methods include random, judgmental, systematic, stratified, and cluster sampling.
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
155
Extraction: Partition and Distribution Coefficients
1.6K
The distribution law or Nernst's distribution law is the law that governs the distribution of a solute between two immiscible solvents. This law, also known as the partition law, states that if a solute is added to the mixture of two immiscible solvents at a constant temperature, the solute is distributed between the two solvents in such a way that the ratio of solute concentrations in the solvents remains constant at equilibrium.
For extracting a solute from an aqueous phase into an...
For extracting a solute from an aqueous phase into an...
1.6K
Aggregates Classification
289
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
289
Cloning of Dolly the Sheep
3.1K
The first successfully cloned mammal was Dolly, a sheep, born on 5th July 1996 at Roslin Institute, Scotland. The cloned sheep was named after the American singer Dolly Parton. Dolly lived for seven years and died of respiratory complications, which is speculated to be due to the actual age of her DNA. Because the DNA in cloned cells belongs to an older individual, the cloned individual’s life expectancy may be affected. Indeed, analysis of Dolly’s DNA revealed shorter...
3.1K
Evolutionary Relationships through Genome Comparisons
5.6K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.6K


