Related Experiment Video
Updated: Oct 3, 2026

Haptic/Graphic Rehabilitation: Integrating a Robot into a Virtual Environment Library and Applying it to Stroke Therapy
Published on: August 8, 2011
AI-augmented ACT-R modeling in real interfaces: automating visual perception, knowledge acquisition, and motor
Amirreza Bagherzadeh1, Farnaz Tehranchi2
1Department of Industrial and Manufacturing Engineering, The Pennsylvania State University, University Park, PA, United States.
Introduction:
Cognitive models in interactive environments often rely on hard-coded symbolic task descriptions, pre-specified interface objects, and idealized action execution. This paper's primary contribution is EVisiTor, an AI-augmented extension of the VisiTor eyes-and-hands tool for ACT-R that automates two forms of work commonly performed by the modeler: perception of the live interface and acquisition of task knowledge.
Methods:
Its visual-perception pathway converts interactive interface objects into ACT-R-compatible visicon features, while its knowledge-acquisition pathway converts human-readable instructions into executable declarative structures. A motor-execution pathway connects ACT-R actions to the operating system and supports probabilistic misexecution with perceptually grounded checking and correction. We use a previously studied Excel procedure as a behavioral case study.
Results:
Relative to an idealized perfect-user model, the EVisiTor-enabled model reduced task-level RMSE from 29.19 s to 17.37 s, human-SD-normalized RMSE from 1.24 to 0.66, and mean absolute per-subtask error from 23.1 s to 10.5 s, with smaller absolute error on 12 of 14 subtasks. A spreadsheet knowledge-acquisition analysis showed that prompting the LLMs with the instruction and UIA grounding can produce the intended action and live visual-object names for every scannable target; the file-name input was the sole exception because it has no stable UIA label. A six-task Windows analysis then tested the feasibility of EVisiTor beyond Excel. Broad instructions often resulted in alternative valid methods: exact action agreement ranged from 42.5%-63.6% without UIA and 18.8%-70.6% with UIA, although 41 of 42 generated parses in each condition were functionally sufficient. Detailed guides raised exact agreement to 70.0%-85.0% without UIA and 76.2%-90.0% with UIA, and validated chunks drove the task-independent model through all six tasks in live Windows applications.
Discussion:
Together, the results demonstrate EVisiTor's potential for automating interface perception and instruction-based knowledge acquisition for a symbolic cognitive model, while identifying instruction specificity and individual variability as targets for future work.