Related Experiment Video
Updated: Mar 7, 2026

Author Spotlight: Integrated Multi-Omics Analysis for Unveiling Multicellular Immune Signatures in Clinical Heart Attack Cohorts
Published on: September 20, 2024
Population-based clustering of co-occurring social determinants: An application of unsupervised machine learning
Ingrid Giesinger1, Emmalin Buajitti1, Arjumand Siddiqi2
1Dalla Lana School of Public Health, University of Toronto, Health Sciences Building, 155 College Street, 6th Floor, Toronto, Ontario M5T 3M7, Canada.
Purpose:
This study aimed to develop a cluster-based measure of multiple co-occurring social determinants of health by applying unsupervised machine learning to a population-based cohort, offering a data-driven approach to organize complex social exposures.
Methods:
Unsupervised clustering was applied to a population-based cohort of Ontario respondents to six-cycles of the Canadian Community Health Survey (2001-2012) linked to the Canadian census and vital statistics data. Clusters were evaluated using internal metrics, visualization techniques, descriptive analysis and theoretical considerations to determine the optimal number of clusters. Sensitivity analyses were integrated across the iterative clustering process. Premature mortality rates were generated assess validity.
Results:
Optimal clustering solutions included 4-clusters and 6-clusters. Both cluster solutions revealed distinct social typologies. The 6-cluster solution offered greater granularity and theoretical interpretability. The 4-cluster solution showed greater heterogeneity within certain marginalized groups. Premature mortality rates differed meaningfully across clusters, supporting the clustering approach in capturing risk associated with social exposure.
Conclusions:
Unsupervised machine learning methods identified meaningful population subgroups reflecting complex patterns of social exposures. This approach offers a flexible, data-driven method for characterizing social exposures that can be considered alongside theoretical frameworks and used for equity monitoring, intervention planning and policy development.
More Related Videos
14:27Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data
Published on: June 26, 2013
12:27Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
Published on: February 15, 2017
Related Concept Videos
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Steps in Outbreak Investigation
Statistical Methods for Analyzing Epidemiological Data
Aggregates Classification
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Sampling Plans
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...