Related Experiment Videos
HAV-Pose: synergizing hybrid attention networks and spherical visibility fields for obstacle-aware fruit harvesting
Yikun Huang1,2, Shuyan Xu3, Yulin Zhong1
1College of Computer and Information Sciences, Fujian Agriculture and Forestry University, Fuzhou, Fujian, China.
Abstract:
The timely removal of crown peppers is a critical agronomic procedure to balance vegetative and reproductive growth. However, automating this task is challenged by severe green-on-green occlusion in dense branching structures and high collision risks within narrow Y-shaped stems. Existing robotic systems suffer from feature erosion when detecting camouflaged targets and lack geometric awareness for safe grasping. To address these challenges, this study proposes Hybrid Attention and Visibility-field based Pose estimation (HAV-Pose), a cascaded perception-planning framework. First, an enhanced detection network based on YOLO11 integrates the original, plug-and-play Hybrid Attention Weighted Convolution (HAWConv) and the RepC3k module, which fuses C3k2 with re-parameterized convolutions to balance inference speed and feature representation. Second, to bridge perception and execution, we propose a Spherical Voxel-based Visibility Field (SVVF) algorithm that transforms complex 3D obstacle avoidance into an efficient 2D visibility search centered on the picking point. SVVF employs a deterministic Max-Margin Optimization Strategy to directly compute collision-free 6D grasping poses with maximal safety margins at millisecond-level efficiency. Extensive experiments evaluate the performance of HAV-Pose. Through benchmarking against 14 SOTA models, structural ablation studies, attention mechanism comparisons, and pixel-wise Euclidean distance evaluations, HAV-Pose achieves a mAP@50 of 90.0% (+8.2%) with high localization precision. In heavily occluded scenarios, SVVF attains a robustness rate of 65.4%, outperforming stochastic baselines (56.8%) with only 67 ms additional computation. Furthermore, generalization is validated across four datasets (Crown Pepper, Green Pepper, Eggplant, and Strawberry) and confirmed through real-world harvesting trials, demonstrating reliable perception-action coupling.