Related Experiment Videos
Adaptive Task Planning for Long-Horizon Robotic Manipulation Based on Video Priors and Dynamic Scene Graphs
Guanghui Ma1, Jiahui Guo2, Xinhua Tang3
1School of Integrated Circuits, Anhui Polytechnic University, Wuhu 241000, China.
Abstract:
Robots are now expected to execute increasingly complex long-horizon tasks in unstructured environments. Despite the strong potential of pretrained Vision-Language Models (VLMs) in task planning, their direct application to robotic manipulation is hindered by logical reasoning deviations and inadequate geometric scene perception. This work proposes an adaptive task planning method based on video priors and dynamic scene graphs (ATP-VPDSG). It leverages the VLM to extract manipulation logic from video demonstrations, thus supplementing manipulation priors. Meanwhile, scene graphs were integrated to convert unstructured environments into structured representations with spatial topological relations, compensating for perceptual deficiencies. A dual-track feedback mechanism based on visual expectations was further incorporated to enable failure diagnosis and adaptive replanning in complex environments. Extensive long-horizon robotic manipulation experiments were conducted on the LIBERO-10 benchmark with Qwen3-VL as the core VLM. Results showed that ATP-VPDSG achieved an average task planning accuracy of 91.2% and a task execution success rate of 74.67%, outperforming the selected task planning baselines. Ablation studies verified that video priors and dynamic scene graphs exerted complementary effects on logical constraints and physical feasibility. Furthermore, a real-robot experiment on an industrial slider-rail assembly task demonstrated successful sim-to-real transfer, achieving an 82.0% success rate without task-specific fine-tuning.