Subgraph Propagation and Contrastive Calibration for Incomplete Multiview Data Clustering
IEEE Transactions on Neural Networks and Learning Systems
|January 18, 2024
Summary
This study introduces a deep clustering framework (SPCC) to address missing data in multiview datasets. SPCC effectively mines incomplete multiview data by reconstructing graph structures and aligning cluster distributions for improved accuracy.
Area of Science:
- Data Mining
- Machine Learning
- Artificial Intelligence
Background:
- Multiview data mining is crucial but challenged by incomplete attributes due to noise and collection failures.
- Existing methods struggle with missing data, failing to calibrate complemented representations with common information across views.
- A significant issue is the cluster distribution unaligned problem (CDUP) in the latent space of incomplete multiview data.
Purpose of the Study:
- To propose a novel deep clustering framework, subgraph propagation and contrastive calibration (SPCC), for incomplete multiview raw data.
- To address the challenges of mining topology in missing multiview data and calibrating complemented representations.
- To solve the cluster distribution unaligned problem (CDUP) in the latent space of incomplete multiview data.
Main Methods:
- Reconstructing a global structural graph by propagating subgraphs from complete data in each view.
- Completing and calibrating missing views guided by the global structural graph and employing contrastive learning between views.
- Aligning complemented cluster distributions across different views using contrastive learning (CL) to resolve CDUP.
Main Results:
- The proposed SPCC framework demonstrates advanced performance on six benchmark datasets.
- The method effectively addresses the challenges of incomplete multiview data, including topology mining and representation calibration.
- Validation confirms the effectiveness and superiority of the SPCC approach in multiview clustering tasks.
Conclusions:
- The SPCC framework offers a robust solution for clustering incomplete multiview data.
- Subgraph propagation and contrastive calibration are effective strategies for handling missing data and aligning latent representations.
- The study validates the SPCC method's ability to achieve superior performance in complex multiview data mining scenarios.
Related Concept Videos
Calibration Curves: Linear Least Squares
1.3K
A calibration curve is a plot of the instrument's response against a series of known concentrations of a substance. This curve is used to set the instrument response levels, using the substance and its concentrations as standards. Alternatively, or additionally, an equation is fitted to the calibration curve plot and subsequently used to calculate the unknown concentrations of other samples reliably.
For data that follow a straight line, the standard method for fitting is the linear...
For data that follow a straight line, the standard method for fitting is the linear...
1.3K
Cluster Sampling Method
11.9K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.9K
Calibration Curves: Correlation Coefficient
1.6K
In a linear calibration curve, there is a value called the calibration coefficient, denoted by 'r,' which measures the strength and the direction of association between two variables. The correlation coefficient value ranges from −1 to +1. A value of +1 indicates a perfect positive linear correlation, −1 denotes a perfect negative correlation, and 0 implies no correlation between the two variables. A positive correlation value establishes that as one variable increases, the...
1.6K
Aggregates Classification
326
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
326
Multicompartment Models: Overview
145
Multicompartment models are mathematical constructs that depict how drugs are distributed and eliminated within the body. They segment the body into several compartments, symbolizing various physiological or anatomical areas connected through drug transfer processes such as absorption, metabolism, distribution, and elimination.
These models offer a more comprehensive representation of drug behavior in the body than one-compartment models. They accommodate the complexity of drug distribution,...
These models offer a more comprehensive representation of drug behavior in the body than one-compartment models. They accommodate the complexity of drug distribution,...
145
Improving Translational Accuracy
10.5K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
10.5K


