Related Experiment Video
Updated: Sep 2, 2025

VisualEyes: A Modular Software System for Oculomotor Experimentation
Published on: March 25, 2011
A Modular Vision Language Navigation and Manipulation Framework for Long Horizon Compositional Tasks in Indoor
Homagni Saha1,2, Fateme Fotouhi1,2, Qisai Liu1
1Department of Mechanical Engineering, Iowa State University, Ames, IA, United States.
We introduce MoViLan, a modular framework for robots to follow household instructions using vision and language. This approach improves performance on complex, long-horizon tasks without needing expert demonstrations.
Area of Science:
- Robotics
- Artificial Intelligence
- Computer Vision
Background:
- Existing end-to-end frameworks struggle with long-horizon, compositional tasks in household settings.
- There's a need for robust methods that handle diverse objects, realistic instructions, and non-reversible state changes.
Purpose of the Study:
- Propose MoViLan (Modular Vision and Language), a novel framework for executing visually grounded natural language instructions for indoor household tasks.
- Address limitations in current approaches for complex manipulation and navigation tasks.
Main Methods:
- Developed a modular approach separating vision and language processing, reducing reliance on strictly aligned training data.
- Introduced a novel geometry-aware mapping technique for cluttered indoor environments.
- Created a generalized language understanding model for household instruction following.
Main Results:
- MoViLan significantly increases success rates for long-horizon, compositional tasks.
- Demonstrated superior performance over recent works on the ALFRED benchmark dataset.
- The modular design allows for more tractable training with separate vision and language datasets.
Conclusions:
- The MoViLan framework offers a more effective and flexible solution for robot instruction following in complex home environments.
- Modular design represents a significant departure from traditional end-to-end methods, enabling better generalization and training efficiency.
More Related Videos
06:28A Networked Desktop Virtual Reality Setup for Decision Science and Navigation Experiments with Multiple Participants
Published on: August 26, 2018
09:46MPI CyberMotion Simulator: Implementation of a Novel Motion Simulator to Investigate Multisensory Path Integration in Three Dimensions
Published on: May 10, 2012
Related Concept Videos
Depth Perception and Spatial Vision
Fluid Movement Between Compartments
Vision
Movement Joints in Buildings
The simplest type of movement joints, working joints, are...
Virtual Work for a System of Connected Rigid Bodies
Next,...
Masonry Curtain Walls