Related Experiment Video
Updated: Jul 7, 2026

05:12
ExCYT: A Graphical User Interface for Streamlining Analysis of High-Dimensional Cytometry Data
Published on: January 16, 2019
Fuzzy c-means clustering of incomplete data
1Math. & Comput. Sci. Dept., Georgia Southern Univ., Statesboro, GA.
Summary
This study introduces four new strategies for fuzzy c-means (FCM) clustering with incomplete data. These methods adapt FCM for datasets with missing feature values, improving clustering accuracy.
Area of Science:
- Data Science
- Machine Learning
- Statistics
Background:
- Clustering algorithms like fuzzy c-means (FCM) are vital for data analysis.
- Standard FCM requires complete data, limiting its application to datasets with missing values.
- Incomplete data is common in real-world scenarios across various scientific domains.
Purpose of the Study:
- To develop and evaluate methods for applying FCM to incomplete datasets.
- To address the challenge of clustering data with missing feature values.
- To enhance the applicability of FCM in practical data mining scenarios.
Main Methods:
- Introduced four distinct strategies for FCM clustering of incomplete data.
- Three strategies involved modified versions of the fuzzy c-means algorithm.
- Numerical convergence properties of the proposed algorithms were analyzed.
Main Results:
- Successfully adapted FCM for datasets containing missing feature values.
- Evaluated the performance of the new algorithms using both real and synthetic incomplete data.
- Demonstrated the feasibility and effectiveness of the proposed clustering strategies.
Conclusions:
- The developed strategies enable effective fuzzy c-means clustering on incomplete datasets.
- These methods expand the utility of FCM for real-world data analysis challenges.
- Further research can explore the convergence and performance in more complex scenarios.
Related Concept Videos
How Data are Classified: Categorical Data
A variable, usually notated by capital letters such as X and Y, is a characteristic or measurement that can be determined for each member of a population. Data are the actual values of variables. They may be numbers, or they may be words. Datum is a single value.
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Cluster Sampling Method
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Aggregates Classification
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Fisher's Exact Test
Fisher's exact test is a statistical significance test widely used to analyze 2x2 contingency tables, particularly in situations where sample sizes are small. Unlike the chi-squared test, which approximates P-values and assumes minimum expected frequencies of at least five in each cell, Fisher's exact test calculates the exact probability (P-value) of observing the data or more extreme results under the null hypothesis. This feature makes it especially valuable when the assumptions of the...
How Data are Classified: Numerical Data
Data that are countable or measurable in specific units are called numerical or quantitative data. Quantitative data are always numbers. Quantitative data are the result of counting or measuring the attributes of a population. Amount of money, pulse rate, weight, number of people living in a town, and number of students who opt for statistics are examples of quantitative data.
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
Flow Cytometry
The development of flow cytometry techniques began in 1934 with initial attempts by Andrew Moldavan, a bacteriologist who counted the cells in a flowing capillary system. Moldavan pumped cells through a capillary tube focused under a microscope for visualization. The invention of photometry allowed the measurement of differentially-stained cells, and Louis Kamentsky developed the first multiparameter flow cytometer in 1965 to identify and count the cancer cells in cervical tissue specimens.
In...
In...
