Related Experiment Video
Updated: Sep 11, 2025

ExCYT: A Graphical User Interface for Streamlining Analysis of High-Dimensional Cytometry Data
Published on: January 16, 2019
Detection and evaluation of clusters within sequential data
Alexander Van Werde1, Albert Senen-Cerda1,2, Gianluca Kosmella1,3
1Department of Mathematics & Computer Science, TU/e, Eindhoven, The Netherlands.
New clustering algorithms for sequential data, based on Block Markov Chains, successfully extract low-dimensional representations from real-world, high-dimensional datasets. These models reveal insights into complex processes like animal movement and DNA sequences.
Area of Science:
- Data Science
- Computational Biology
- Bioinformatics
Background:
- Sequential data is prevalent in various fields, presenting challenges due to high dimensionality, sparsity, and noise.
- Extracting meaningful insights from complex sequential processes requires robust methods to handle data dependencies.
Purpose of the Study:
- To evaluate novel clustering algorithms, derived from Block Markov Chains theory, on real-world sequential data.
- To determine if these algorithms can effectively generate useful low-dimensional representations from sparse, high-dimensional sequences.
Main Methods:
- Application of new clustering algorithms designed for sequential data.
- Empirical study across diverse real-world datasets including animal movement (GPS), DNA sequences, text, and financial data.
- Analysis of the extracted low-dimensional representations for their ability to encode sequential structure and reveal underlying process characteristics.
Main Results:
- The algorithms successfully extracted low-dimensional representations from diverse real-world sequential data.
- These representations effectively captured the inherent sequential structure within the datasets.
- The identified representations provided novel insights into the complex processes under study.
Conclusions:
- The Block Markov Chain-based clustering algorithms are effective for extracting meaningful low-dimensional representations from complex, real-world sequential data.
- This approach offers a promising method for gaining deeper understanding in fields dealing with sequential information.
- The study validates the utility of these algorithms beyond synthetic data, demonstrating their applicability in practical scenarios.
Related Concept Videos
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Sampling Plans
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Steps in Outbreak Investigation
Detection of Gross Error: The Q Test

