Related Experiment Videos
Cross-Attention-Driven Propose-and-Select Generative Data Augmentation for Few-Shot Image Classification
Ying Liu1,2, Liaomo Zheng1,3, Shiyu Wang1,3
1Shenyang Institute of Computing Technology, Chinese Academy of Sciences, Shenyang 110168, China.
Abstract:
Generative data augmentation based on diffusion models has emerged as a promising approach for few-shot image classification. Existing methods, such as DA-Fusion, typically follow a "generate-once, use-directly" paradigm, which often suffers from uncontrollable generation quality, unstable semantic consistency, insufficient global diversity, and high sample redundancy. To address these limitations, we propose a two-stage Propose-and-Select framework for controllable data augmentation. This framework curates high-quality synthetic data offline, ensuring that no additional training overhead is introduced to downstream models. For selector optimization, our method eliminates the need for additional human annotations by leveraging the zero-shot prior knowledge of a vision-language model (CLIP) to construct relative-quality pseudo-labels. Furthermore, we develop an adaptive-temperature listwise ranking distillation objective to transfer quality-aware supervision effectively. We also introduce a multi-objective consistency regularization strategy to stabilize training and improve convergence. Under a strictly controlled augmentation budget, where all methods are provided with the same number of synthetic samples, the proposed approach consistently outperforms existing diffusion-based augmentation baselines across both few-shot classification benchmarks, achieving an accuracy of 79.58% on PASCAL VOC and 80.74% on the fine-grained Oxford 102 Flowers dataset. These results demonstrate the effectiveness of the proposed generation-selection paradigm in improving the quality, diversity, and semantic relevance of synthetic samples, thereby enhancing downstream few-shot classification performance.