Related Experiment Videos
Target-Aware Decoupled Metric-Depth Estimation for Top-View Crane Safety Surveillance
Min Woo Woo1,2, Jaeil Kim1, Byeong Hak Kim2
1School of Computer Science and Engineering, IT College, Kyungpook National University, Daegu 41566, Republic of Korea.
Abstract:
Estimating metric depth for safety-relevant targets from a top-view crane camera is challenging because monocular predictions are scale-ambiguous and target regions of interest (RoIs) mix returns from payloads, personnel, hoisting wires, and the ground. We present a target-focused pipeline in which a detector defines RoIs, a dense metric-depth map is predicted, and one representative depth is extracted per object. A DINOv2-DPT relative-depth network is fixed, while a multiscale adapter, DINO detection branch, and pixel-wise spatial affine calibration head are trained. Ground-truth bounding boxes restrict metric supervision to target-relevant regions but are not inputs to the calibration head. At evaluation, all depth methods use identical predicted RoIs, and a class-aware kernel-density estimation rule extracts representative depths. On 272 hardware-synchronized RGB-LiDAR frames from one operational 150-ton crane, the proposed method achieved a target-level RMSE of 0.564 m and a pixel-level inlier ratio of 0.966 for δ<1.25. One-to-one ground-truth matching yielded a 0.566 m target RMSE with detection precision/recall of 0.960/0.970. AdaBins obtained lower pixel-level MAE and RMSE, whereas the proposed method obtained lower target-level errors. These results demonstrate the feasibility of target-focused metric ranging under the evaluated operational crane conditions.