Robustness assessment of regressions using cluster analysis typologies: a bootstrap procedure with application in
Leonard Roth1, Matthias Studer2, Emilie Zuercher3
1Department of Epidemiology and Health Systems, Centre for Primary Care and Public Health (Unisanté), University of Lausanne, Route de La Corniche 10, 1010, Lausanne, Switzerland. leonard.roth@unisante.ch.
BMC Medical Research Methodology
|December 18, 2024
Summary
Standard sequence analysis may yield incorrect conclusions by ignoring sampling uncertainty. This study introduces a robust method to assess regression results, ensuring reliable findings in trajectory pattern analysis.
Area of Science:
- Social Sciences
- Biomedical Data Science
Background:
- Standard sequence analysis clusters trajectories but often ignores sampling uncertainty.
- This oversight can lead to inaccurate conclusions in regression models linking patterns to covariates.
Purpose of the Study:
- To introduce a novel procedure for assessing the robustness of regression results in sequence analysis.
- To account for sampling uncertainty in trajectory typology and associated regressions.
Main Methods:
- Utilized bootstrap samples to construct new typologies and estimate regression models.
- Employed a multilevel modeling framework, mimicking meta-analysis, to combine bootstrap estimates.
- Applied the methodology to healthcare utilization trajectories in a Swiss diabetic patient cohort.
Main Results:
- The procedure yields robust estimates and 95% prediction intervals, accounting for sampling uncertainty.
- Identified central and borderline trajectories within clusters.
- In the illustrative application, the association between lipid testing and healthcare utilization was not supported upon robustness assessment.
Conclusions:
- The developed Robustness Assessment of Regression using Cluster Analysis Typologies (RARCAT) enhances the reliability of association studies.
- RARCAT is applicable to any analysis combining clustering with regression, extending beyond state sequence analysis.
Related Concept Videos
Survival Tree
58
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
58
Bootstrapping
583
The term "bootstrap" originated in the 19th century as a metaphor for self-improvement or achieving something independently, without external assistance. This concept extends to statistical bootstrapping, a self-contained method for estimating population parameters through resampling, even though it can be computationally intensive. Developed by the American statistician Dr. Bradley Efron in 1979, bootstrapping provides a robust way to perform inference when the original sample size is...
583
Cluster Sampling Method
11.6K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.6K
Comparing the Survival Analysis of Two or More Groups
149
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
149
Variability: Analysis
126
Measures of variability are statistical metrics that reveal the dispersion pattern within a dataset. They are pivotal in biostatistics, providing insights into the heterogeneity within health and biological data. Variability signifies the degree to which data points diverge from one another, helping researchers understand the potential range of values and associated uncertainty within the data.
The range is a simple measure of variability, indicating the difference between the highest and...
The range is a simple measure of variability, indicating the difference between the highest and...
126
Quantifying and Rejecting Outliers: The Grubbs Test
1.5K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.5K


