Related Experiment Video
Updated: Oct 7, 2026

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
HAE-Net: A hierarchy-aware enhanced method for 3D panoptic segmentation in orchard scenes
Shuo Zhang1,2, Yiqun Wang1,2, Lianchuan Shi1,2
1School of Computer Science and Artificial Intelligence, Shandong Normal University, Jinan, 250358, China.
Abstract:
Reliable perception of orchard scenes is needed for agricultural automation. Fruit detection, counting, and localization support yield prediction, harvest planning, and precision fertilization, while 3D point clouds provide geometric information with less dependence on ambient illumination than conventional RGB images. Orchard analysis also requires predictions at more than one scale: individual fruits and trunks must be separated, but the fruits must still be associated with their parent tree. Most existing 3D panoptic segmentation methods operate at a single instance granularity and do not represent this hierarchy. We propose HAE-Net, a hierarchy-aware segmentation method built on a shared encoder and multiple decoders. At the encoder output, the Hierarchical Feature Enhancement (HFE) module supplies global aggregation for tree-level grouping, multi-scale context for semantic classification, and local boundary information for standard-instance separation. The Cross-Decoder Attention Gate (CDAG) regulates the encoder features transferred at each skip connection. Its spatial, channel, multi-scale, and pixel-level components allow the three decoders to use different mixtures of the shared representation. A Hierarchical Consistency Loss (HCL) links the two instance granularities by constraining the predicted centers of component instances and their parent tree instance. We evaluate HAE-Net on HOPS (Hierarchical Orchard Panoptic Segmentation), which includes point clouds acquired with terrestrial laser scanners, uncrewed aerial vehicles, uncrewed ground vehicles, and SfM reconstruction. Quantitative comparisons and ablations use the labeled TLS validation split. HAE-Net achieves an mIoU of 71.63% and an mPQ of 53.96%, exceeding HAPT3D by 9.81 and 5.43 percentage points, respectively. Because ground-truth labels for the four test subsets are not publicly accessible, their predictions are used only for qualitative inspection.
Related Concept Videos
Depth Perception and Spatial Vision
Parallel Processing