Tensor Low-Rank Representation for Data Recovery and Clustering.
IEEE Transactions on Pattern Analysis and Machine Intelligence
|November 22, 2019
Summary
This study introduces Tensor Low-Rank Representation (TLRR), a novel method for analyzing tensor data. TLRR accurately recovers corrupted data and clusters it effectively, offering provable performance guarantees for various applications.
Area of Science:
- Multi-way data analysis
- Tensor decomposition
- Machine learning
Background:
- Tensor data analysis is increasingly important in various fields.
- Existing methods struggle with corrupted data and accurate clustering.
- Low-rank representation is a promising approach for data analysis.
Purpose of the Study:
- To develop a Tensor Low-Rank Representation (TLRR) method.
- To achieve exact recovery of clean tensor data with intrinsic low-rank structure.
- To accurately cluster tensor data with provable performance guarantees.
Main Methods:
- Developed a novel Tensor Low-Rank Representation (TLRR) method.
- Proposed efficient convex programming for optimizing the TLRR objective function.
- Introduced two dictionary construction methods: simple TLRR (S-TLRR) and robust TLRR (R-TLRR).
Main Results:
- TLRR exactly recovers clean tensor data from arbitrary sparse corruptions under mild conditions.
- TLRR accurately verifies true origin tensor subspaces for precise clustering.
- Experimental results show superior performance, efficiency, and robustness over state-of-the-art methods.
Conclusions:
- TLRR is the first method to exactly recover and accurately cluster tensor data with intrinsic low-rank structure.
- TLRR offers provable performance guarantees and is optimized via efficient convex programming.
- S-TLRR and R-TLRR effectively handle slightly and severely corrupted data, respectively.
Related Concept Videos
Cluster Sampling Method
13.9K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
13.9K
Reducing Line Loss
331
In a three-phase circuit, line loss is an indicator of energy dissipated as heat due to the resistance of transmission lines. To address this, incorporating transformers into the system—a step-up transformer at the source and a step-down transformer at the load—is a strategic solution. Two three-phase transformers are introduced to improve this.
With a step-up transformer at the source, the voltage is increased, thereby reducing the current in the transmission lines since power loss in...
With a step-up transformer at the source, the voltage is increased, thereby reducing the current in the transmission lines since power loss in...
331
Residuals and Least-Squares Property
8.9K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
8.9K
Extraction: Partition and Distribution Coefficients
4.5K
The distribution law or Nernst's distribution law is the law that governs the distribution of a solute between two immiscible solvents. This law, also known as the partition law, states that if a solute is added to the mixture of two immiscible solvents at a constant temperature, the solute is distributed between the two solvents in such a way that the ratio of solute concentrations in the solvents remains constant at equilibrium.
For extracting a solute from an aqueous phase into an...
For extracting a solute from an aqueous phase into an...
4.5K
Regression Toward the Mean
6.8K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.8K
Survival Tree
356
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
356


