Related Experiment Video
Updated: Jan 15, 2026

Methods to Explore the Influence of Top-down Visual Processes on Motor Behavior
Published on: April 16, 2014
VLA-MP: A Vision-Language-Action Framework for Multimodal Perception and Physics-Constrained Action Generation in
Maoning Ge1, Kento Ohtani1, Yingjie Niu1
1Graduate School of Informatics, Nagoya University, Furo-cho, Chikusa-Ward, Nagoya 464-8601, Japan.
Abstract:
Autonomous driving in complex real-world environments requires robust perception, reasoning, and physically feasible planning, which remain challenging for current end-to-end approaches. This paper introduces VLA-MP, a unified vision-language-action framework that integrates multimodal Bird's-Eye View (BEV) perception, vision-language alignment, and a GRU-bicycle dynamics cascade adapter for physics-informed action generation. The system constructs structured environmental representations from RGB images and LiDAR, aligns scene features with natural language instructions through a cross-modal projector and large language model, and converts high-level semantic hidden states outputs into executable and physically consistent trajectories. Experiments on the LMDrive dataset and CARLA simulator demonstrate that VLA-MP achieves high performance across the LangAuto benchmark series, with best driving scores of 44.3, 63.5, and 78.4 on LangAuto, LangAuto-Short, and LangAuto-Tiny, respectively, while maintaining high infraction scores of 0.89-0.95, outperforming recent VLA methods such as LMDrive and AD-H. Visualization and video results further validate the framework's ability to follow complex language-conditioned instructions, adapt to dynamic environments, and prioritize safety. These findings highlight the potential of combining multimodal perception, language reasoning, and physics-aware adapters for robust and interpretable autonomous driving.
Related Concept Videos
Depth Perception and Spatial Vision
Three-Dimensional Force System:Problem Solving
To solve a three-dimensional force system, first resolve each force into its respective scalar components. Do this using...
Inertial Frames of Reference
Muscle Coordination and Action
Agonists
Agonist muscles, often called prime movers, are the primary muscles responsible for producing a specific movement....
Hierarchy of Motor Control
Non-inertial Frames of Reference

