Related Experiment Video
Updated: Jan 9, 2026

Eye Tracking During Visually Situated Language Comprehension: Flexibility and Limitations in Uncovering Visual Context Effects
Published on: November 30, 2018
Proposal of a Focused Self-Attention LLM Allowing to Interpret Fragmentary Requests from Visually Impaired People
Abstract:
The development of Artificial Intelligence (AI) has been widely applied as an assistive technology for individuals with visual impairment (VI), providing various assistive technologies such as object recognition and obstacle detection. The accurate and powerful systems can be triggered through voice commands and thus significantly enhance the life experiences of users with VI. However, most assistive technologies simply accept a clear and exact command as input, ignoring the cognition and memory problems in blind people. If the users forget some parts of the full command (e.g., name of a target object, available system actions), it will be difficult for them to utilize the assistive technologies. To fill the research gap, our study presents an interactive system which integrates a Large Language Model (LLM) with focused self-attention (FSA) mechanisms to assistive technologies, facilitating the voice control for VI users. A virtual reality-based experimental study (N=12) was conducted to simulate a real-world application scenario to compare the direct voice command control with the proposed LLM-FSA approach. The results demonstrated significant improvements in convenience, intuitiveness, and efficiency when using the LLM-FSA method. This research provides insights into the integration of domain-specific LLMs into assistive technologies, contributing to more accessible solutions for visually impaired individuals.

