MAGIC: Meta-Ability Guided Interactive Chain-of-Distillation for Effective-and-Efficient Vision-and-Language

Summary

This study introduces the MAGIC method for creating smaller, efficient embodied artificial intelligence (E-AI) models for Vision-and-Language Navigation (VLN) tasks. The proposed approach significantly outperforms existing methods, even with drastically reduced model sizes.