DR-CLIP: A Deformable Vision-Language Model for Scale-Invariant Object Counting in Remote Sensing Images

Jingzhe Nie1, Qun Liu1, Tianze Li1

  • 1College of Information Science and Engineering, Shandong Agricultural University, Tai'an 271018, China.

Summary

DR-CLIP enhances remote sensing object counting using deformable attention and text guidance. This vision-language model improves accuracy for small and dense objects, enabling open-vocabulary counting without retraining.