基于能源的聚类:以已知的概率函数进行快速和强大的数据聚类
Moritz Thürlemann1, Sereina Riniker1
1Department of Chemistry and Applied Biosciences, ETH Zürich, Vladimir-Prelog-Weg 2, 8093 Zürich, Switzerland.
The Journal of chemical physics
|July 10, 2023
概括
基于能源的集群 (EBC) 提供了一种新的方法来分析复杂的数据,减少对密度估计的依赖. 这种以分子动力学为灵感的方法提高了聚类的准确性和效率.
科学领域:
- 计算化学是一种计算化学.
- 数据科学是数据科学.
- 统计力学就是统计力学.
背景情况:
- 聚类算法通常依赖于密度估计,这很容易受到大数据集中的维度和采样问题的诅咒.
- 分子动力学模拟产生复杂的数据,其中密度估计可能不可靠.
研究的目的:
- 开发一种新的集群算法,即基于能源的集群 (EBC),以减轻对估计数据密度的依赖.
- 使用大都会的接受标准来概括光谱聚类,并纳入潜在能源信息.
主要方法:
- 开发了一个基于能源的集群 (EBC) 算法,利用大都会的接受标准.
- 制定了EBC作为在高温下光谱聚类的概括.
- 将潜在能量直接纳入聚类过程,允许对密集区域进行分样.
主要成果:
- 通过利用潜在能源信息,EBC有效地将集群与采样密度脱.
- 由于高效的子采样,该算法展示了显著的加速度和亚线性缩放.
- 验证了阿拉宁二和Trp子小蛋白的分子动力学轨迹的EBC.
结论:
- 基于能源的聚类为复杂数据集提供了密度依赖方法的强大替代方案.
- 潜在能量的直接包含提高了模拟中的集群性能和计算效率.
- EBC为分析大规模分子动力学和其他复杂数据提供了一个有希望的方向.
相关概念视频
Cluster Sampling Method
12.0K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
12.0K
Expected Frequencies in Goodness-of-Fit Tests
2.6K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.6K
Probability Histograms
11.8K
A probability histogram is a visual representation of a probability distribution. Similar a typical histogram, the probability histogram consists of contiguous (adjoining) boxes. It has both a horizontal axis and a vertical axis. The horizontal axis is labeled with what the data represents. The vertical axis is labeled with probability. Each rectangular bar in the histogram is 1 unit wide, which suggests that the area under each bar equals the probability, P(x), where x is 1, 2, 3, and so on.
11.8K
Determination of Expected Frequency
2.2K
Suppose one wants to test independence between the two variables of a contingency table. The values in the table constitute the observed frequencies of the dataset. But how does one determine the expected frequency of the dataset? One of the important assumptions is that the two variables are independent, which means the variables do not influence each other. For independent variables, the statistical probability of any event involving both variables is calculated by multiplying the individual...
2.2K
Probability Distributions
7.3K
The probability of a random variable x is the likelihood of its occurrence. A probability distribution represents the probabilities of a random variable using a formula, graph, or table. There are two types of probability distribution– discrete probability distribution and continuous probability distribution.
A discrete probability distribution is a probability distribution of discrete random variables. It can be categorized into binomial probability distribution and Poisson...
A discrete probability distribution is a probability distribution of discrete random variables. It can be categorized into binomial probability distribution and Poisson...
7.3K
Probability in Statistics
13.5K
Probability is the likelihood of an event occurring. The term event is defined as a collection of results of a procedure. An event is a simple event when an outcome cannot be divided into simpler parts.
An example of a simple event is a coin toss. The result of a coin toss is either a head or a tail. Here, head and tail are two simple events. These two simple events make up the sample space. Further, the probability of an event occurring falls within the range of 0 to 1. The probability of an...
An example of a simple event is a coin toss. The result of a coin toss is either a head or a tail. Here, head and tail are two simple events. These two simple events make up the sample space. Further, the probability of an event occurring falls within the range of 0 to 1. The probability of an...
13.5K


