MIFuzzy Clustering for Incomplete Longitudinal Data in Smart Health
1Department of Computer and Information Science, University of Massachusetts Dartmouth, Department of Quantitative Health Sciences, University of Massachusetts Medical School, Worcester, MA 01655.
Smart Health (Amsterdam, Netherlands)
|October 11, 2017
Summary
Multiple imputation (MI) enhances fuzzy clustering for smart health studies with missing data. This method improves clustering accuracy, offering better insights into treatment effect heterogeneity.
Area of Science:
- Data Science
- Machine Learning
- Health Informatics
Background:
- Missing data frequently occur in longitudinal smart health studies, impacting data analysis.
- Multiple imputation (MI) is a standard technique for handling missing data but its application in unsupervised learning needs further investigation.
- Fuzzy clustering is a valuable tool for analyzing complex datasets, including those with missing values.
Purpose of the Study:
- To introduce and evaluate the MIFuzzy clustering approach for handling missing data in smart health studies.
- To theoretically, empirically, and numerically demonstrate the performance of MI-based fuzzy clustering.
- To compare the uncertainty of clustering accuracy between MI-based and traditional imputation methods.
Main Methods:
- Developed and applied the MIFuzzy clustering algorithm, an extension of previous MIfuzzy clustering work.
- Utilized non-parametric soft computing techniques for semi-supervised and unsupervised learning.
- Compared MIFuzzy clustering against non-imputation and single-imputation based clustering approaches.
Main Results:
- MI-based fuzzy clustering significantly reduces uncertainty in clustering accuracy compared to other methods.
- The MIFuzzy approach demonstrates robust performance in processing incomplete longitudinal behavioral intervention data.
- Theoretical, empirical, and numerical analyses confirm the utility and strength of the proposed method.
Conclusions:
- MIFuzzy clustering offers a powerful solution for analyzing incomplete longitudinal data in smart health.
- This approach enhances the reliability of clustering results, crucial for understanding treatment effect heterogeneity.
- The study advances the application of imputation techniques within unsupervised learning frameworks for health research.
Related Concept Videos
Cluster Sampling Method
15.0K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
15.0K
Statistical Methods for Analyzing Epidemiological Data
1.0K
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
1.0K
Longitudinal Studies
563
Longitudinal studies are also widely used in other medical and social science fields. For instance, in cardiovascular research, they can monitor patients' health over decades to identify risk factors for heart disease, such as high cholesterol or smoking, and evaluate the long-term effectiveness of preventive measures. Similarly, in mental health studies, researchers might follow individuals from adolescence into adulthood to understand the development and progression of conditions like...
563
Longitudinal Research
13.5K
Sometimes we want to see how people change over time, as in studies of human development and lifespan. When we test the same group of individuals repeatedly over an extended period of time, we are conducting longitudinal research. Longitudinal research is a research design in which data-gathering is administered repeatedly over an extended period of time. For example, we may survey a group of individuals about their dietary habits at age 20, retest them a decade later at age 30, and then again...
13.5K
Classification of Illness
9.0K
The meaning of illness is individualized to each person who experiences an alteration in health. In contrast, disease is a medical term indicating a pathological change in the structure and function of the body or mind. It is a condition that has specific symptoms and boundaries.
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
9.0K
Biostatistics: Overview
927
Biostatistics plays a crucial role in understanding and analyzing data in healthcare and biology. Biostatisticians conduct experiments, gather evidence, and draw meaningful conclusions using statistical methods and techniques. Different variables form the foundation of biostatistical analysis, allowing researchers to understand and interpret data effectively. These variables are classified into different types, each serving a specific purpose in statistical analysis.
Discrete variables are...
Discrete variables are...
927


