Somtimes:用于时间序列聚类的自我组织地图及其对严重疾病对话的应用
Ali Javed1,2, Donna M Rizzo3,2, Byung Suk Lee2
1Department of Medicine, Stanford University, 300 Pasteur Dr, Stanford, CA 94305 USA.
概括
一个新的算法,SOMTimeS,通过使用动态时间曲 (DTW) 提高速度和可扩展性,高效地集群大型时间序列数据. 这种方法增强了对复杂数据集的时间序列分析.
科学领域:
- 人工智能的人工智能
- 机器学习 机器学习
- 数据科学数据科学数据科学
背景情况:
- 需要可扩展的算法来聚类和分析大时间序列数据.
- 科霍宁的自组织地图 (SOM) 是有效的集群和缩小维度.
- 动态时间扭曲 (DTW) 准确地测量时间序列相似性,但由于二次复杂性,它在计算上昂贵.
研究的目的:
- 介绍SOMTimeS,一个使用DTW.W的时间序列集群的新型自组织地图.
- 为了提高基于DTW的时间序列集群的可扩展性和速度.
- 评估SOMTimeS的性能与现有方法相比.
主要方法:
- 开发了SOMTimeS,这是一个Kohonen自组织地图,使用动态时间扭曲 (DTW) 作为距离测量.
- 实施了修剪策略,以减少在SOM培训期间不必要的DTW计算.
- 引入了K-TimeS,这是一个K-means变体,具有类似的修剪策略进行比较.
主要成果:
- SOMTimeS显示了与其他基于DTW的聚类算法相似的准确性.
- 修剪策略显著提高了计算性能,将DTW计算减少了高达50%.
- 对112个基准数据集的评估显示,SOMTimeS和K-TimeS的平均速度是1.8倍,性能因数据集而异.
结论:
- SOMTimeS提供了一种准确且显著更快的方法来对大型时间序列数据集进行聚类.
- 修剪技术有效地解决了DTW的计算限制.
- 在实践中,SOMTimeS具有实用性,包括在医疗保健环境中对自然语言分析的应用.
相关概念视频
Classification of Illness
7.5K
The meaning of illness is individualized to each person who experiences an alteration in health. In contrast, disease is a medical term indicating a pathological change in the structure and function of the body or mind. It is a condition that has specific symptoms and boundaries.
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
7.5K
Time-Series Graph
4.4K
A time-series graph is a line graph with repeated measurements taken at successive intervals of time. It is also called a time series chart. To construct a time-series graph, one must look at both pieces of a paired data set. The horizontal axis is used to plot the time increments, and the vertical axis is used to plot the values of the variable that one is measuring. By using the axes in this way, each point on the graph will correspond to time and a measured quantity. The points on the graph...
4.4K
Cluster Sampling Method
11.9K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.9K
Causality in Epidemiology
395
Causality or causation is a fundamental concept in epidemiology, vital for understanding the relationships between various factors and health outcomes. Despite its importance, there's no single, universally accepted definition of causality within the discipline. Drawing from a systematic review, causality in epidemiology encompasses several definitions, including production, necessary and sufficient, sufficient-component, counterfactual, and probabilistic models. Each has its strengths and...
395


