Related Experiment Video
Updated: Jul 15, 2026

Measuring Attention and Visual Processing Speed by Model-based Analysis of Temporal-order Judgments
Published on: January 23, 2017
Learning robust and generalizable bimanual skills: a spatiotemporal causal hierarchical diffusion framework with
Xukun Liu1, Fengjuan Xie1, Zhenyu Liu1
1Northwest Institute of Mechanical and Electrical Engineering, Xianyang, Shaanxi, China.
Introduction:
Bimanual visuomotor imitation learning enables robots to acquire coordinated dual-arm manipulation skills from visual demonstrations, yet it faces significant challenges in temporal synchronization, spatial collision avoidance, long-horizon reasoning, and robustness to visual distractions. Existing diffusion-based policies often struggle to simultaneously capture long-horizon temporal dependencies and fine-grained spatial precision, while remaining sensitive to spurious correlations and domain shifts.
Methods:
To address these limitations, we propose the Spatiotemporal Causal Hierarchical Diffusion Imitation Learner (SCH-DIL), a framework that integrates spatiotemporal hierarchical diffusion optimization to factorize the denoising process into temporal and spatial branches for multi-scale action modeling. The framework further incorporates causal visual representation learning that minimizes mutual information with confounding environmental factors to produce invariant features, along with noise-robust diffusion modeling that employs learnable observation uncertainty estimation and confidence-aware denoising. Additionally, attention anti-interference regularization is introduced to penalize distractions and enforce temporal attention consistency.
Results:
Extensive experiments on the RoboTwin 2.0 benchmark demonstrate that SCH-DIL consistently outperforms existing diffusion-based and imitation learning baselines, achieving higher success rates under both clean and domain-randomized inference settings.
Discussion:
These improvements are achieved with minimal computational overhead, suggesting that the proposed hierarchical and causally regularized diffusion framework offers a practical and robust solution for bimanual visuomotor imitation learning in visually challenging and domain-shifted environments.