関連する実験動画
Updated: Jan 17, 2026

Decoding Natural Behavior from Neuroethological Embedding
Published on: October 3, 2025
アンサンブルエンコーダーによるマルチモーダルおよび柔軟なスケールデータを用いたプロアクティブな人間組立意図認識
Abstract:
Human-robot collaboration (HRC) assembly necessitates precise mutual cognition to guarantee safe and efficient execution. In this context, human assembly intention recognition (HAIR) serves as a critical approach to achieving this mutual understanding. However, most current HAIR approaches struggle to extract sufficient spatiotemporal information from limited industrial data, particularly under complex conditions like varying scales and visual occlusions. Thereby, this article proposes an ensemble encoder approach to extract and fuse spatial and temporal features from visual and skeleton streams of the HRC assembly process, thus significantly improving HAIR accuracy and efficiency. First, an RGB feature extraction encoder is designed to model spatiotemporal dependencies of the assembly process with different scales of features from flexible input RGB encoders (RGBEs). Distinctively, a cross-attention module is utilized to fuse information from different-scale RGBEs, ensuring comprehensive assembly action representation with different granularities. Second, to address the occlusion challenge, a mask-aware skeleton feature extraction encoder is devised. By utilizing frame and joint masking strategies, it robustly models the relationship between operator pose evolution and assembly actions, maintaining high performance even under occlusion. Third, a global feature fusion encoder integrates and aligns features from RGB and skeleton feature extraction encoders. Experimental results demonstrate the state-of-the-art performance of the proposed approach, which achieves the highest accuracy of 99.12%, 99.23%, and 84.59% on MCV-Intention, HA4M, and HA-VID datasets, respectively. Six ablation studies demonstrate the performance effects of fusion positions, the number of depth channels, cross-attention fusion module, occlusions, illuminations, and computational efficiency.
関連する概念動画
Automatic Processing and Automatic Social Behavior
Multi-input and Multi-variable systems
In the absence of...

