Related Experiment Videos
An Augmented Reality and AI-Based System for Contextual Appliance Guidance: Implications for Cognitive Accessibility
Kimia Hafezi1, Atra Hossein Tafreshi1, Christian Napoli1,2
1Department of Computer, Control, and Management Engineering "Antonio Ruberti", Sapienza University of Rome, 00185 Rome, Italy.
Background:
Modern household appliances often present complex interfaces that can be difficult to use, especially for older adults, people with visual impairments, and users with mild cognitive difficulties. In such cases, interacting with appliances may require sustained attention, visuospatial search, working memory, and sequential action planning, while traditional user manuals often provide limited contextual support.
Methods:
To address this issue, this study presents a proof-of-concept augmented reality (AR) and artificial intelligence (AI)-based system for contextual appliance guidance. The proposed architecture integrates visual sensing, deep learning, and large language models to detect appliance controls, interpret user queries, retrieve relevant information from user manuals, and provide step-by-step guidance directly on the real interface. A YOLOv8 model trained on a custom dataset was used for button detection, YOLO-Seg was employed to enhance visual highlighting through segmentation, and BoT-SORT was used to maintain detection consistency across frames. A Unity-based mobile application displayed real-time AR overlays with customizable visual settings for accessibility needs, such as low vision and color blindness that may be relevant for future accessibility-oriented applications. In addition to its technical pipeline, the system is conceptually relevant as a potential form of external cognitive support because it transforms static manual instructions into situated, sequential, and visually grounded guidance.
Results:
Experimental results showed promising technical performance for button detection and segmentation, while a preliminary user evaluation in a non-clinical sample suggested good usability, clarity, and acceptability of the interface.
Conclusions:
These findings support the technical feasibility and preliminary usability of the approach, while cognitive workload, confidence, functional autonomy, and clinical benefit were not directly measured. Targeted validation in older adults, people with visual impairments, and clinical populations is therefore still needed. Future developments will include multimodal feedback, read-aloud guidance, and more specific evaluation of workload, confidence, and functional autonomy.