Related Experiment Video
Updated: Feb 14, 2026

A Practical Guide to Phylogenetics for Nonexperts
Published on: February 5, 2014
Use of clustering techniques for clinical and epidemiological research: practical tips using an example from
Martine Jansen1, Mark M Bakker2, Alexandre R Sepriano3
1Center for Statistics (CenStat), Data Science Institute (DSI), Hasselt University, Diepenbeek, Belgium; Marketing, Communication & PR, Fontys University of Applied Sciences, Eindhoven, The Netherlands.
None:
Clinical and epidemiological researchers may want to identify or determine groups with similar characteristics that exist within a dataset or population. For this purpose, clustering techniques can be applied. However, clustering can be complex, and results can be misleading if the methods are not applied with fidelity. The aim of this paper is to provide practical guidance to (clinical) researchers in rheumatology and other fields who wish to explore or confirm (sub) groups within their study population, and are considering to apply clustering techniques to (patient) data. We discuss when clustering is useful for addressing your research question (and when it is not), the need to define the cluster concept, the choice of distance measure between 2 observations, the preprocessing of the data, the assessment of clustering tendency, the types of clustering methods, the determination of the number of clusters and end with the evaluation of the derived clustering solutions using visualizations and statistical measures. The following methods are presented in this paper: K-means, partitioning around medoids, hierarchical clustering, Density-Based Spatial Clustering of Applications with Noise, spectral clustering, fuzzy clustering, and latent class analysis. To illuminate the different considerations involved in clustering, we use a case example: secondary data from 895 people with rheumatoid arthritis, spondyloarthritis, and gout who completed the Health Literacy Questionnaire (HLQ). In this case example, we compute, appraise and evaluate clustering solutions using different methods in 6 steps, including statistical measures and expert opinion. The 6 steps proposed in this paper describe a systematic approach to cluster analysis to ensure all important aspects are considered. In the case example, different clustering methods led to different clustering solutions and insights. Qualitative interpretation by content experts was insightful here. Where possible, researchers should apply more than one clustering method fitting to the objective and reflect on eventual differences in solutions.
More Related Videos
10:11Fundus Photography as a Convenient Tool to Study Microvascular Responses to Cardiovascular Disease Risk Factors in Epidemiological Studies
Published on: October 22, 2014
08:55A Practical Guide for the Production and PET/CT Imaging of 68Ga-DOTATATE for Neuroendocrine Tumors in Daily Clinical Practice
Published on: April 17, 2019
Related Concept Videos
Introduction to Epidemiology
Causality in Epidemiology
Study Designs in Epidemiology
Observational studies are those where the researcher does not intervene but rather observes natural variations. They include cross-sectional, cohort, and...
Confounding in Epidemiological Studies
Bias in Epidemiological Studies
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...