A supervised clustering MCMC methodology for large categorical feature spaces
Simón Ramírez1, Adolfo J Quiroz2, Alvaro J Riascos3
1University of Califonia, Berkeley, United States and Quantil, Bogotá, Colombia.
Statistical Methods in Medical Research
|June 2, 2021
Summary
This study introduces a supervised clustering method to reduce large feature spaces, improving machine learning model accuracy. Applied to health insurance, it enhances risk adjustment by grouping diagnostic codes for better expenditure prediction.
Area of Science:
- Statistics
- Machine Learning
- Health Informatics
Background:
- Dimensionality reduction is crucial in machine learning, often addressed via unsupervised methods.
- Large categorical feature spaces pose challenges for predictive modeling.
- Existing risk adjustment models in health insurance can be improved.
Purpose of the Study:
- To introduce a novel supervised clustering methodology for optimizing large categorical feature spaces.
- To enhance the accuracy of supervised learning algorithms by minimizing test error.
- To improve risk adjustment in health insurance markets through better health expenditure prediction.
Main Methods:
- Developed a supervised clustering methodology utilizing a Metropolis Hastings algorithm.
- Optimized the partition structure of large categorical feature spaces.
- Applied the methodology to cluster International Classification of Diseases, 10th Revision (ICD-10) codes.
Main Results:
- The supervised clustering approach demonstrated superior performance compared to common alternatives.
- Clustering diagnostic codes into risk groups improved health expenditure prediction.
- The methodology was validated on a large dataset from the Colombian Healthcare System.
Conclusions:
- The proposed supervised clustering method effectively reduces dimensionality and minimizes test error.
- This approach offers a significant improvement for risk adjustment in health insurance.
- The methodology has broad applicability to supervised learning problems with large categorical features.
Related Concept Videos
Cluster Sampling Method
13.4K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
13.4K
Sampling Plans
470
Sampling is a crucial step in analytical chemistry, allowing researchers to collect representative data from a large population. Common sampling methods include random, judgmental, systematic, stratified, and cluster sampling.
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
470
How Data are Classified: Categorical Data
39.0K
A variable, usually notated by capital letters such as X and Y, is a characteristic or measurement that can be determined for each member of a population. Data are the actual values of variables. They may be numbers, or they may be words. Datum is a single value.
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
39.0K
Survival Tree
201
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
201
Vesicular Tubular Clusters
2.7K
After budding out from the ER membrane, some COPII vesicles lose their coat and fuse with one another to form larger vesicles and interconnected tubules called vesicular tubular clusters or VTCs. These clusters constitute a compartment at the ER-Golgi interface known as ERGIC (Endoplasmic Reticulum Golgi Intermediate Compartment). The ERGIC is a mobile membrane-bound cargo transport system that sorts proteins secreted from ER and delivers them to the Golgi.
With the help of motor proteins such...
With the help of motor proteins such...
2.7K
Collisions in Multiple Dimensions: Problem Solving
4.6K
In multiple dimensions, the conservation of momentum applies in each direction independently. Hence, to solve collisions in multiple dimensions, we should write down the momentum conservation in each direction separately. To help understand collisions in multiple dimensions, consider an example.
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
A small car of mass 1,200 kg traveling east at 60 km/h collides at an intersection with a truck of mass 3,000 kg traveling due north at 40 km/h. The two vehicles are locked together. What is the...
4.6K


