Related Experiment Video
Updated: Jun 14, 2025

07:11
Author Spotlight: Insights into Visual Cortex Research Through Wide-View fMRI Mapping
Published on: December 8, 2023
1.4K
ECCT: Efficient Contrastive Clustering via Pseudo-Siamese Vision Transformer and Multi-view Augmentation
Xing Wei1, Taizhang Hu2, Di Wu3
1School of Computer and Information, Hefei University of Technology, Hefei, 230601, China; Intelligent Manufacturing Technology Research Institute, Hefei University of Technology, Hefei, 230601, China.
Summary
This study introduces Efficient Contrastive Clustering via Pseudo-Siamese Vision Transformer and Multi-view Augmentation (ECCT) for improved image clustering. ECCT enhances feature representation by integrating global information and semantic continuity, outperforming existing methods.
Area of Science:
- Computer Vision
- Machine Learning
- Deep Learning
Background:
- Image clustering is crucial for organizing unlabeled image data.
- Contrastive learning methods excel at learning discriminative features but struggle with global information and semantic continuity.
- Existing methods often have limited feature distributions, hindering contrastive learning's potential in clustering.
Purpose of the Study:
- To propose a novel deep clustering framework, Efficient Contrastive Clustering via Pseudo-Siamese Vision Transformer and Multi-view Augmentation (ECCT).
- To address the limitations of existing methods in capturing global information, preserving semantic continuity, and broadening feature distributions.
- To enhance the performance of image clustering using advanced deep learning techniques.
Main Methods:
- Introduced a pseudo-Siamese Vision Transformer (ViT) framework incorporating a Hilbert Patch Embedding (HPE) module for global feature extraction.
- Fused features from two ViT branches to achieve both global view and semantic coherence.
- Employed multi-view random aggressive augmentation to diversify feature distributions and learn richer contrastive features.
Main Results:
- ECCT demonstrated superior performance across five benchmark datasets compared to existing clustering methods.
- Achieved an Adjusted Rand Index (ARI) of 0.852 on the STL-10 dataset, outperforming the best baseline by 10.3%.
- On the ImageNet-Dogs dataset, ECCT reached an ARI of 0.424, an improvement of 4.8% over the best baseline.
Conclusions:
- ECCT effectively addresses the challenges of global information capture and semantic continuity in image clustering.
- The proposed framework significantly enhances the learning of comprehensive and rich contrastive features.
- ECCT represents a substantial advancement in deep clustering, offering improved accuracy and robustness.

