Related Experiment Video
Updated: Sep 16, 2025

Development of an Audio-based Virtual Gaming Environment to Assist with Navigation Skills in the Blind
Published on: March 27, 2013
Egocentric Perception of Walking Environments Using an Interactive Vision-Language System
Abstract:
Large language models can provide detailed contextual understanding of a scene beyond what computer vision alone can provide, which has implications for robotics and embodied intelligence. In this study, we developed a novel multi-modal vision-language system for egocentric visual perception, with an initial focus on real-world walking environments. We trained a number of state-of-the-art transformer-based vision-language models that use causal language modelling on our custom dataset of 43,055 image-text pairs for few-shot image captioning. We designed a new speech synthesis model and a user interface to convert the generated image captions into speech for audio feedback to users. Our system also uniquely allows for feedforward user prompts to personalize the generated image captions by incorporating human cognition in the decision making. Our system is able to generate detailed captions with an average length of 10 words while achieving a high ROUGE-L score of 43.9% and a low word error rate of 28.1% with an end-to-end processing time of 2.2 seconds. Overall, our new multimodal vision-language system can generate accurate and detailed descriptions of scenes, which can be further augmented by user prompts. This innovative feature allows our image captions to be personalized to the individual needs and preferences of the user, thus optimizing the closed-loop interactions between the human and generative AI models for understanding and navigating of real-world environments.
Related Concept Videos
Depth Perception and Spatial Vision
Perception
Bottom-up processing begins at the sensory level, where receptors detect external environmental stimuli. These could include the tactile sensation of...
Visual System
Once through the pupil, the light passes through the lens, a...
Gestalt Principles of Perception
Vision

