Related Experiment Video
Updated: Aug 9, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving
Summary
Researchers developed DriveMonkey, a framework enhancing large vision-language models (LVLMs) for autonomous driving. It improves 3D spatial understanding and instruction following, outperforming existing models in 3D visual grounding.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Robotics
Background:
- Large Vision-Language Models (LVLMs) show promise for autonomous driving but struggle with comprehensive scene understanding.
- Existing LVLMs lack 2D-3D mapping and integrated 3D spatial reasoning with instruction following.
- Current research often uses limited scene data and simple annotations, hindering complex task performance.
Purpose of the Study:
- To address limitations in LVLMs for autonomous driving by improving 3D spatial understanding and instruction following.
- To introduce a novel dataset and framework for comprehensive scene understanding and interactive tasks.
- To enhance language-conditioned 3D grounding capabilities in complex driving environments.
Main Methods:
- Introduced NuInteract, a large-scale dataset with over 1.5M multi-view image-language pairs for dense scene captions and interactive tasks.
- Proposed DriveMonkey, a framework integrating LVLMs with a spatial processor via learnable queries.
- Utilized a plug-and-play spatial processor, initialized with 3D detectors, for structured geometric priors in 3D grounding.
Main Results:
- DriveMonkey demonstrates superior performance compared to general LVLMs in autonomous driving scenarios.
- Achieved a significant 9.86% improvement on the challenging 3D visual grounding task.
- The proposed framework effectively integrates 3D spatial understanding with language-conditioned instruction following.
Conclusions:
- DriveMonkey offers a significant advancement in enabling LVLMs for complex autonomous driving tasks.
- The NuInteract dataset and DriveMonkey framework provide valuable resources for future research in vision-language understanding and robotics.
- Future work can leverage these resources to develop more robust and capable AI systems for real-world applications.