Related Experiment Video
Updated: Jul 16, 2026

An Assessment Method and Toolkit to Evaluate Keyboard Design on Smartphones
Published on: October 5, 2020
Word-Level Motion Learning for Contactless QWERTY Typing with a Single Camera
Sung-Sic Yoo1, Heung-Shik Lee1
1Department of Smart Mobility Engineering, Joongbu University, 305 Dongheon-ro, Deogyang-gu, Goyang-si 21713, Gyeonggi-do, Republic of Korea.
None:
Contactless text entry is increasingly important in immersive and constrained computing environments, yet most vision-based approaches rely on character-level recognition or key localization, which are fragile under monocular sensing. This study investigates the feasibility of recognizing natural QWERTY typing motions directly at the word level using only a single RGB camera, under a fixed single-user and single-camera configuration. We propose a word-level contactless typing framework that models each word as a distinctive spatiotemporal finger motion pattern derived from hand joint trajectories. Typing motions are temporally segmented, and direction-aware finger displacements are accumulated to construct compact motion representations that are relatively insensitive to absolute hand position and typing duration within the evaluated setup. Each word is represented by multiple motion prototypes that are incrementally updated through online learning with a trial-delayed adaptation protocol. Experiments with vocabularies of up to 200 words show that the proposed approach progressively learns and recalls word-level motion patterns through repeated interaction, achieving stable recognition performance within the tested configuration at realistic typing speeds. Additional evaluations demonstrate that learned motion representations can transfer from physical keyboards to flat-surface typing within the same experimental setting, even when tactile feedback and visual layout cues are reduced. These results support the feasibility of reframing contactless typing as a word-level motion recall problem, and suggest its potential role as a complementary component to character-centric camera-based input methods under constrained monocular sensing.

