Related Experiment Videos
Multimodal graph-based fusion via image descriptions for few-shot open-set recognition
1School of Information Technology and Engineering, Guangzhou College of Commerce, Guangzhou, 511363, China.
Abstract:
Few-shot open-set recognition (FSOR) presents unique challenges due to the limited labeled samples and the presence of unknown classes during inference. While recent few-shot learning methods leverage class-level textual information to enhance feature discrimination, they often overlook image-grounded descriptions that provide more fine-grained and instance-specific semantics. Moreover, many of these approaches require prior knowledge of the class name from labeled samples during inference, which may be impractical in open-world scenarios. In this study, we propose a multimodal Graph-based Fusion (MGF) framework that learns visually grounded semantic representations to enhance FSOR performance. MGF leverages image descriptions generated by a vision-language model as the textual supervision to guide the learning of semantic features from images. A graph convolutional network is then used to fuse semantic and visual features, enabling effective intra-class information propagation and improving discrimination between known and unknown classes. We jointly optimize a contrastive and a semantic alignment loss to promote intra-class compactness and inter-class separability. Extensive experiments on several few-shot learning benchmarks demonstrate that MGF achieves superior open-set recognition and competitive closed-set classification performance.