Related Experiment Video
Updated: Apr 18, 2026

Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data
Published on: June 26, 2013
A general framework for a reliable multivariate analysis and pattern recognition in high-dimensional epidemiological
1Inserm, UMR_S 1136, department of social epidemiology, Pierre-Louis institute of epidemiology and public health, 27, rue de Chaligny, 75012 Paris, France; UMR_S 1136, UPMC université Paris 06, Sorbonne universités, 75646 Paris, France; Inserm, UMR_S 1101, laboratory of medical information processing, 29609 Brest, France; UMR_S 1101, université de Bretagne occidentale, 29609 Brest, France; UMR_S 1101, institut Mines-Telecom, Telecom Bretagne, 29609 Brest, France.
Epidemiologists can enhance their analysis with clustering techniques, powerful multivariate tools for pattern recognition and group discovery without prior assumptions. These methods offer a robust alternative to traditional statistical approaches for identifying homogeneous groups in data.
Area of Science:
- Epidemiology
- Biostatistics
- Data Science
Background:
- Classical epidemiological tools include comparisons of means/proportions, regression models (linear, logistic, Cox), and their multivariate formulations.
- Natively massive multivariate techniques, such as clustering, are underutilized in epidemiology despite weaker assumptions and pattern recognition capabilities.
- Clustering is widely applied in genetics and biomolecular studies for identifying homogeneous groups without a priori knowledge.
Purpose of the Study:
- To introduce epidemiologists to underutilized multivariate clustering techniques.
- To explain the core principles of clustering, including proximity definition and optimal group number determination.
- To present archetypal algorithms like k-means and PAM and their relation to classical statistical methods.
Main Methods:
- Clustering techniques require parameter tuning, notably the number of groups to discover.
- Approaches like the silhouette and robustness methods aid in finding the optimal number of groups.
- The study details key aspects of clustering: defining observation proximity and determining the number of groups, using k-means and PAM as examples.
Main Results:
- A framework is presented to reconsider epidemiological concerns using clustering.
- The study demonstrates how to identify group existence, determine the optimal number of groups, and label observations.
- Methods for analyzing identified groups with explicative data and achieving consistent, initialization-insensitive results are shown.
Conclusions:
- Clustering techniques, combined with parameter tuning, offer substantial new tools for epidemiologists.
- Unlike hypothesis-testing approaches, clustering methods make no assumptions on data and are natively multivariate.
- These techniques enable pattern recognition and homogeneous group retrieval, expanding the epidemiologist's analytical capabilities.
Related Concept Videos
Statistical Methods for Analyzing Epidemiological Data
Bias in Epidemiological Studies
Biostatistics: Overview
Discrete variables are...
Confounding in Epidemiological Studies
Statistical Software for Data Analysis and Clinical Trials
Statistical Analysis: Overview
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...

