Consensus Kernel K-Means Clustering for Incomplete Multiview Data
Yongkai Ye1, Xinwang Liu1, Qiang Liu1
1College of Computer, National University of Defense Technology, Changsha, China.
Computational Intelligence and Neuroscience
|January 10, 2018
Summary
This study introduces a novel method for incomplete multiview clustering that imputes missing data and ensures consistency across views. The approach enhances clustering performance by explicitly modeling between-view consistency for better results.
Area of Science:
- Machine Learning
- Data Science
- Computer Vision
Background:
- Multiview clustering integrates information from multiple data sources to improve clustering accuracy.
- Existing methods struggle with incomplete data across different views.
- Prior work unified multiview clustering and imputation but neglected between-view consistency.
Purpose of the Study:
- To propose a unified learning framework for incomplete multiview clustering.
- To address the limitations of existing methods by incorporating between-view consistency.
- To simultaneously impute missing data and learn a consistent clustering result.
Main Methods:
- A novel unified learning method is proposed for incomplete multiview clustering.
- The method explicitly models between-view consistency by measuring similarity between individual view clusters and a consensus cluster.
- Incomplete views are imputed to optimize clustering within each view while preserving inter-view consistency.
Main Results:
- The proposed method demonstrates superior performance on both synthetic and real-world incomplete multiview datasets.
- Experimental comparisons show significant improvements over state-of-the-art methods.
- The explicit modeling of between-view consistency is shown to be crucial for enhanced performance.
Conclusions:
- The proposed unified learning method effectively handles incomplete multiview clustering by integrating imputation and consistency modeling.
- The approach offers a significant advancement in multiview clustering research, particularly for datasets with missing information.
- The findings highlight the importance of inter-view consistency for robust multiview clustering.
More Related Videos
Related Concept Videos
Incomplete Dominance
30.2K
Gregor Mendel's work (1822 - 1884) was primarily focused on pea plants. Through his initial experiments, he determined that every gene in a diploid cell has two variants called alleles inherited from each parent. He suggested that amongst these two alleles, one allele is dominant in character and the other recessive. The combination of alleles determines the phenotype of a gene in an organism.
30.2K
Cluster Sampling Method
14.9K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
14.9K
Vesicular Tubular Clusters
3.3K
After budding out from the ER membrane, some COPII vesicles lose their coat and fuse with one another to form larger vesicles and interconnected tubules called vesicular tubular clusters or VTCs. These clusters constitute a compartment at the ER-Golgi interface known as ERGIC (Endoplasmic Reticulum Golgi Intermediate Compartment). The ERGIC is a mobile membrane-bound cargo transport system that sorts proteins secreted from ER and delivers them to the Golgi.
With the help of motor proteins such...
With the help of motor proteins such...
3.3K
How Data are Classified: Categorical Data
45.5K
A variable, usually notated by capital letters such as X and Y, is a characteristic or measurement that can be determined for each member of a population. Data are the actual values of variables. They may be numbers, or they may be words. Datum is a single value.
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
45.5K
How Data are Classified: Numerical Data
38.6K
Data that are countable or measurable in specific units are called numerical or quantitative data. Quantitative data are always numbers. Quantitative data are the result of counting or measuring the attributes of a population. Amount of money, pulse rate, weight, number of people living in a town, and number of students who opt for statistics are examples of quantitative data.
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
38.6K
Data Reporting and Recording
5.5K
Reporting and recording are crucial in data documentation. The timely, thorough, and accurate documentation of facts is essential when recording patient data. Failure to record findings during an assessment or interpretation of a problem will result in loss of information and make the patient document unreliable. The reader is left with general impressions if the information is not specific. A recording is documenting data of the individual's health information in a traceable, secure, and...
5.5K


