CLUSTERnGO: a user-defined modelling platform for two-stage clustering of time-series data
Işık Barış Fidaner1, Ayca Cankorur-Cetinkaya2, Duygu Dikicioglu2
1Department of Computer Engineering.
Bioinformatics (Oxford, England)
|September 29, 2015
Summary
CLUSTERnGO is a new bioinformatics tool for analyzing time-series data, excelling with transient phenomena. It uses advanced clustering to reveal novel biological insights and outperforms existing methods.
Area of Science:
- Bioinformatics
- Computational Biology
- Statistical Analysis
Background:
- Standard bioinformatics tools struggle with transient phenomena in time-series data.
- This limitation hinders the extraction of meaningful biological information.
- There is a need for specialized, user-friendly tools for time-series analysis.
Purpose of the Study:
- To introduce CLUSTERnGO, a novel statistical application for time-series data analysis.
- To address the limitations of existing tools in handling transient phenomena.
- To provide a flexible and user-friendly solution for biological data interpretation.
Main Methods:
- Development of CLUSTERnGO, a model-based clustering application.
- Utilizes an Infinite Mixture of Piecewise Linear Sequences (Bayesian non-parametric model).
- Employs a novel Two-Stage Clustering methodology and Gene Ontology (GO) enrichment analysis with multiple hypothesis testing.
Main Results:
- CLUSTERnGO successfully analyzes time-series datasets, including transient phenomena.
- The application outperforms existing algorithms in assigning unique GO term enrichments.
- Novel biological insights were uncovered in diverse test cases, surpassing original publication findings.
Conclusions:
- CLUSTERnGO offers a powerful and flexible approach to time-series data analysis.
- The tool enhances biological interpretation through GO enrichment.
- It provides valuable insights for both specialist and non-specialist users.
Related Concept Videos
Time-Series Graph
5.6K
A time-series graph is a line graph with repeated measurements taken at successive intervals of time. It is also called a time series chart. To construct a time-series graph, one must look at both pieces of a paired data set. The horizontal axis is used to plot the time increments, and the vertical axis is used to plot the values of the variable that one is measuring. By using the axes in this way, each point on the graph will correspond to time and a measured quantity. The points on the graph...
5.6K
Cluster Sampling Method
15.6K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
15.6K
Survival Tree
496
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
496
Comparing the Survival Analysis of Two or More Groups
710
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
710
Sampling Plans
1.3K
Sampling is a crucial step in analytical chemistry, allowing researchers to collect representative data from a large population. Common sampling methods include random, judgmental, systematic, stratified, and cluster sampling.
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
1.3K
Friedman Two-way Analysis of Variance by Ranks
567
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
567


