Related Experiment Video
Updated: Mar 29, 2026

08:25
Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
9.7K
DR-CLIP: A Deformable Vision-Language Model for Scale-Invariant Object Counting in Remote Sensing Images
Jingzhe Nie1, Qun Liu1, Tianze Li1
1College of Information Science and Engineering, Shandong Agricultural University, Tai'an 271018, China.
Sensors (Basel, Switzerland)
|March 28, 2026
Summary
DR-CLIP enhances remote sensing object counting using deformable attention and text guidance. This vision-language model improves accuracy for small and dense objects, enabling open-vocabulary counting without retraining.
Area of Science:
- Computer Vision
- Remote Sensing
- Artificial Intelligence
Background:
- Object counting in remote sensing is crucial for urban planning and environmental monitoring.
- Challenges include varied annotations, ambiguous queries, and poor small target detection.
- Existing methods struggle with open-vocabulary queries and scale variations.
Purpose of the Study:
- To develop a robust vision-language model for accurate and open-vocabulary object counting in remote sensing images.
- To address limitations of heterogeneous annotations and small object detection.
- To improve cross-modal retrieval performance in remote sensing contexts.
Main Methods:
- Proposed DR-CLIP (Deformable Remote CLIP) incorporating deformable visual feature extraction and text-guided prediction.
- Introduced a Region-to-Instruction (R2I) mechanism for unified annotation representation.
- Implemented Multi-scale Deformable Attention (MSDA) for enhanced feature extraction and a Text-Guided Counting Head for cross-modal alignment.
Main Results:
- Achieved Mean Absolute Error (MAE) of 2.34 and Root Mean Squared Error (RMSE) of 3.89 on DOTA-v2.0, outperforming baselines by 19.0% in MAE.
- MSDA module improved Small-Object Recall (SOR) to 0.824, effective for dense and small object counting.
- Attained R@1 scores of 68.3% (image-to-text) and 72.1% (text-to-image) on RSICD dataset.
Conclusions:
- DR-CLIP offers a powerful solution for open-vocabulary object counting in remote sensing, overcoming key challenges.
- The model demonstrates superior performance, particularly in detecting small and dense objects.
- DR-CLIP exhibits robust generalization capabilities across different domains with minimal performance degradation.

