Related Experiment Video
Updated: Sep 27, 2026

Methods to Explore the Influence of Top-down Visual Processes on Motor Behavior
Published on: April 16, 2014
Restoring the Right Stream: Training-Free OOD Robustness for Vision-Language-Action Policies
1Smart Tower Co., Ltd., Beijing 100195, China.
Abstract:
Vision-Language-Action (VLA) policies remain brittle under modest distribution shift. On LIBERO-Plus, contemporary models that solve clean tasks at high rates can fall below 30% success when the camera's viewpoint or the robot's initial pose is perturbed. Most training-free test-time remedies address this problem through the image stream, for example, by augmenting, purifying, or selecting visual observations. In our controlled evaluation, this family of methods improves mean success by only about three points and leaves the robot-initial-state failure largely unresolved. This paper studies the failure at the level of input streams. A VLA receives visual tokens, a proprioceptive state token, and language tokens; different perturbations can move different streams away from their training manifold. In particular, the robot-initial-state perturbation directly shifts the proprioceptive token; therefore, image-space interventions have limited leverage. We introduce Gated Per-Stream Manifold Restoration (G-PSMR), a training-free wrapper for a frozen policy. For each stream, a lightweight gate detects off-manifold inputs and applies a stream-specific restoration before the policy forward pass. We instantiate the framework with entropy-gated visual consensus and gated relative-orientation debiasing, which preserves the within-episode orientation trajectory. In the original 280-episode paired evaluation, the joint method improves total success by +5.3 points compared with a +3.2 image-only gain and raises the most fragile factor from 20% to 38%. On 1120 previously unevaluated, manifest-disjoint task instances, the state restoration improves robot-initial-state success from 23.8% to 28.7%; a separate prospectively specified confirmation on 600 new gate-active instances yields 26.3%→32.2% (+5.8 points; 95% CI [+3.2,+8.5]; p<0.001). Together, the original cross-stream results and two independent state-stream evaluations support the central principle of matching the restoration to the input stream carrying the shift.
Related Concept Videos
Fixed Action Patterns
Vision
Observational Learning
Woodward–Hoffmann Selection Rules and Microscopic Reversibility
Purposive Learning
Introduction to Learning
In contrast to learned behaviors, unlearned behaviors such as crying, sexual...