Related Experiment Video
Updated: May 31, 2026

ExCYT: A Graphical User Interface for Streamlining Analysis of High-Dimensional Cytometry Data
Published on: January 16, 2019
Clustering methods for categorical time series and sequences : a scoping review
Ottavio Khalifa1, Alan Balendran2, Viet-Thi Tran2,3
1Université Paris Cité, Université Sorbonne Paris Nord, INSERM, INRAE, Centre for Research in Epidemiology and StatisticS (CRESS), Paris, France. ottavio.khalifa@inserm.fr.
This review overviews clustering methods for categorical time series (CTS), common in many fields. A new typology and web tool aid researchers in selecting appropriate CTS clustering techniques based on data characteristics.
Area of Science:
- Data Science
- Bioinformatics
- Epidemiology
- Sociology
- Marketing Analytics
Background:
- Categorical time series (CTS) data are prevalent across diverse scientific disciplines, including epidemiology, sociology, biology, and marketing.
- Clustering these complex data structures presents unique challenges, such as variable sequence lengths, multivariate observations, and missing data.
Purpose of the Study:
- To systematically review and categorize existing clustering methods specifically designed for categorical time series (CTS) data.
- To provide a framework for selecting appropriate CTS clustering methods based on specific data characteristics and research objectives.
Main Methods:
- A comprehensive literature search was conducted across PubMed, Web of Science, and Google Scholar up to November 2024.
- Identified CTS clustering methods were classified into three main families: distance-based, feature-based, and model-based.
- Methods were evaluated based on their ability to handle common challenges like variable sequence length, multivariate data, and missing values.
Main Results:
- Out of 14,607 retrieved records, 124 articles describing 129 distinct CTS clustering methods were included in the review.
- Distance-based methods, particularly those employing Optimal Matching, were the most frequently proposed (56 methods).
- Model-based methods (28) demonstrated greater capacity for complex data structures, while feature-based methods (45) offered better scalability at the expense of flexibility.
- Public implementations were available for fewer than half of the reviewed methods.
- A searchable web application was developed to facilitate method selection.
Conclusions:
- Categorical time series clustering methods exhibit significant heterogeneity in their underlying assumptions, capabilities, and scalability.
- While distance-based methods are prevalent, model-based approaches offer more sophisticated modeling potential for complex data.
- The developed typology and web application serve as valuable resources for researchers navigating the diverse landscape of CTS clustering techniques.
Related Concept Videos
Time-Series Graph
Comparing the Survival Analysis of Two or More Groups
Review and Preview
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Survival Tree
Building a Survival Tree
Constructing a survival tree begins...
