RAR: Retrieving and Ranking Augmented MLLMs for Visual Recognition

Summary

This study introduces RAR, a novel method combining CLIP and Multimodal Large Language Models (MLLMs) to improve fine-grained visual recognition. RAR enhances few-shot and zero-shot capabilities for extensive, detailed datasets.

Related Concept Videos