Related Experiment Video
Updated: Jun 30, 2026

Simultaneous Long-term Recordings at Two Neuronal Processing Stages in Behaving Honeybees
Published on: July 21, 2014
ST-HONet: Spatio-Temporal Hierarchical Network for long-horizon bimanual visuomotor imitation
Xukun Liu1, Fengjuan Xie1, Kai Xu1
1Northwest Institute of Mechanical and Electrical Engineering, Xianyang, Shaanxi, China.
Introduction:
Learning robust and temporally consistent manipulation policies from long-horizon visual observations remains a fundamental challenge in imitation learning. While recent Transformer-based approaches reduce compounding errors via temporally extended action chunks, most methods rely on deterministic or unimodal action representations, limiting their ability to capture the inherent multimodality of expert demonstrations.
Methods:
In this work, we propose ST-HONet, a Spatio-Temporal Hierarchical Network that integrates three adaptive optimization mechanisms operating jointly across spatial and temporal dimensions. ST-HONet formulates policy learning as conditional generation of extended action segments using a Transformer-based conditional variational autoencoder to explicitly model multimodal expert behaviors through a structured latent representation. To ensure stable optimization over long temporal horizons, we introduce a unified spatio-temporal training framework combining adaptive data augmentation, progressive latent regularization, and multi-stage optimization strategies.
Results:
We evaluate ST-HONet on RoboTwin 2.0, a large-scale benchmark for long-horizon bimanual manipulation under domain randomization with automatically generated expert demonstrations. Across representative manipulation tasks, ST-HONet achieves consistently higher task success rates compared to baseline models, while incurring only minimal additional computational overhead.
Discussion:
These results demonstrate that explicitly modeling multimodality via a structured latent space, combined with joint spatio-temporal training mechanisms, significantly improves policy robustness and temporal consistency in long-horizon visuomotor imitation learning. The minimal computational overhead further supports the practical deployability of ST-HONet in real-world systems.
Related Concept Videos
Nonconscious Mimicry
Observational Learning
Hierarchy of Motor Control
Facial Feedback Hypothesis

