Related Experiment Video
Updated: Jun 21, 2025

Heuristic Mining of Hierarchical Genotypes and Accessory Genome Loci in Bacterial Populations
Published on: December 7, 2021
Cluster Analysis Methods to Support Population Health Improvement Among US Counties
Elizabeth A Pollock1, Ronald E Gangnon, Keith P Gennuso
1Department of Population Health Sciences, University of Wisconsin Population Health Institute, University of Wisconsin-Madison, Madison, Wisconsin (Drs Pollock, Gennuso and Givens); and Department of Population Health Sciences, University of Wisconsin-Madison, Madison, Wisconsin (Dr Gangnon).
Context:
Population health rankings can be a catalyst for the improvement of health by drawing attention to areas in need of relative improvement and summarizing complex information in a manner understood by almost everyone. However, ranks also have unintended consequences, such as being interpreted as "hard truths," where variations may not be significant. There is a need to improve communication about uncertainty in ranks, with accurate interpretation. The most common solutions discussed in the literature have included modeling approaches to minimize statistical noise or borrow strength from covariates. However, the use of complex models can limit communication and implementation, especially for broad audiences.
Objectives:
Explore data-informed grouping (cluster analysis) as an easier-to-understand, empirical technique to account for rank imprecision that can be effectively communicated both numerically and visually.
Design:
Cluster analysis, specifically k-means clustering with Wasserstein (earth mover's) distance, was explored as an approach to identify natural and meaningful groupings and gaps in the data distribution for the County Health Rankings' (CHR) health outcomes ranks.
Setting:
County-level health outcomes from the 2022 CHR.
Participants:
3082 counties that were ranked in the 2022 CHR.
Main Outcome Measure:
Data-informed health groups.
Results:
Cluster analysis identified 30 health groupings among counties nationwide, with cluster size ranging from 9 to 184 counties. On average, states had 16 identified clusters, ranging from 3 in Delaware and Hawaii to 27 in Virginia. Number of clusters per state was associated with number of counties per state and population of the state. The method helped address many of the issues that arise from providing rank estimates alone.
Conclusions:
Public health practitioners can use this information to understand uncertainty in ranks, visualize distances between county ranks, have context around which counties are not meaningfully different from one another, and compare county performance to peer counties.
More Related Videos
06:55Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
11:21Methodology for Establishing a Community-Wide Life Laboratory for Capturing Unobtrusive and Continuous Remote Activity and Health Data
Published on: July 27, 2018
Related Concept Videos
Statistical Methods for Analyzing Epidemiological Data
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Methods Of Healthcare Delivery System
Managed Care System:
The managed care system is designed to control the cost while maintaining the quality of care. The patient's care from admission to discharge is planned by the primary care provider or the case manager, also known as the gatekeeper. In a managed care system, the number of care providers is...
Steps in Outbreak Investigation
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Sampling Plans
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...