Related Experiment Video
Updated: May 28, 2026

Artificial Intelligence-Based System for Detecting Attention Levels in Students
Published on: December 15, 2023
Harmonizing Scale for Intelligent Sensors: Resolution-Conditioned Adaptation of Vision Foundation Models for Crowd
Huan Xu1, Zhiheng Chen1, Sirou Shen1
1School of Computer Engineering, Jimei University, Xiamen 361021, China.
None:
Scale variation remains a fundamental challenge for intelligent surveillance sensors in crowd-counting and localization tasks. While recent selective inheritance methods have shown promise through multi-resolution feature fusion, they typically rely on conventional CNN backbones with limited representation capacity. Foundation models such as DINOv3 offer powerful self-supervised representations, yet directly applying frozen DINOv3 features to selective scale-aware counting is non-trivial: these features are semantically strong but not directly aligned with density estimation or scale-conditioned inheritance requirements. In this paper, we propose D3-CalibCount, a trainable-parameter-efficient framework for adapting frozen DINOv3 representations to selective scale-aware crowd counting and localization. We introduce a lightweight Scale Harmonization Adapter (SHA) that performs resolution-conditioned feature calibration, transforming generic DINOv3 representations into scale-selective counting features suitable for progressive inheritance across resolution levels. Extensive experiments on three widely used benchmarks show consistent improvements over selective inheritance baselines, especially under severe scale variation. From a deployment perspective, the method reduces the trainable portion of the network, while the frozen DINOv3-L backbone still introduces a higher inference cost than lighter CNN baselines. The target application scenario is camera-based intelligent surveillance sensing, where crowd density estimation is performed from visual sensor inputs at GPU-backed monitoring nodes. These results suggest that lightweight adaptation of frozen foundation features is a practical direction for crowd counting and other dense prediction tasks.
