Related Experiment Video
Updated: May 27, 2026

Quantification of Oculomotor Responses and Accommodation Through Instrumentation and Analysis Toolboxes
Published on: March 3, 2023
Edge-Based Vision-Language Assistive System for the Visually Impaired: A Quantized VLM Approach
None:
This paper presents a novel edge-deployable assistive system for visually impaired individuals, powered by vision-language models (VLMs). Existing cloud-based solutions for image captioning suffer from latency, dependency on internet connectivity, and overly simplistic scene descriptions that fail to convey the rich contextual information needed for meaningful real-world navigation. To address these challenges, we propose a self-contained, wearable system that performs real-time scene interpretation on-device without cloud reliance. The system integrates a quantized version of the LLaVA-NeXT-13B VLM with speech recognition (Whisper) and speech synthesis (PIPER), running on an NVIDIA Jetson Orin NX and Raspberry Pi-based input module. Our framework emphasizes intuitive user interaction through a button-based interface and Bluetooth audio output, minimizing cognitive load. We validate the quantization approach through large-scale benchmarks, such as VizWiz-VQA and VQAv2, demonstrating minimal accuracy degradation (2.6% and 0.9%, respectively) compared to the original model. User evaluations involving 28 participants compared the system to a baseline image captioning model. Objective results demonstrated a 25% increase in image identification accuracy. The system achieved high usability scores and maintained practical latency (4-5 seconds per query), supporting real-world feasibility. This work advances the development of scalable, interpretable, and accessible AI-driven assistive technologies, enabling greater independence and interaction for the visually impaired.

