Related Experiment Video
Updated: Mar 21, 2026

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Large-scale training data enhances silent speech decoding with around-ear EEG
Masakazu Inoue1, Eri Hatakeyama1, Yuya Kita1
1Sanpo Sakuma Building 6 F 1-11 Kanda Sakuma-cho, Chiyoda-ku, Tokyo 101-0025, Japan.
None:
Objective. Silent speech decoding (SSD) offers a potential communication alternative for individuals with impaired vocalization. However, conventional multi-electrode electroencephalography (EEG) or facial electromyography (EMG) systems require cumbersome preparation and are unsuitable for daily use. This study evaluates the practicality of SSD using a wearable around-ear EEG device, focusing on data scaling, cross-subject transfer, vocabulary extensibility, and online decoding performance.Approach. We collected 72 h of around-ear EEG from 24 healthy participants and one individual with incomplete locked-in syndrome (LIS) during silent, vocalized, and attempted speech, and integrated these around-ear EEG recordings with prior EMG + high-density EEG datasets, yielding 282.4 total h of training data. Using a 64-word classification task as the evaluation metric, we assessed: (1) whether larger datasets improve around-ear EEG-based SSD, (2) whether healthy-participant data supplement limited LIS-participant data despite articulatory differences, (3) transferability to unseen vocabulary, and (4) online user-interface performance.Main results. Large-scale EEG/EMG data improved SSD accuracy in both healthy participants and the LIS participant. Training on the heterogeneous dataset achieved 56.6% accuracy for healthy users and 47.3% for the LIS participant. Fine-tuning this decoder for new vocabulary increased the accuracy by 22 percentage points relative to training from scratch. Regression analysis showed that, for decoding in the LIS participant, data from the LIS participant contributed approximately four times the weight of healthy-participant data, quantifying data strategies for SSD. Online experiments achieved top-1/top-5 accuracies of 47.2%/76.0% for healthy users and 26.5%/49.1% for the LIS participant.Significance. The results indicate that lightweight, commercially feasible around-ear EEG can enable practical SSD when combined with large-scale healthy-participant data, supporting online operation. Moreover, models trained on a 64-word vocabulary facilitate decoding of a new vocabulary, providing a path toward SSD systems requiring minimal LIS-participant data. This study advances non-invasive SSD systems suitable for everyday communication.
More Related Videos
11:39Assessment of Audio-Tactile Sensory Substitution Training in Participants with Profound Deafness Using the Event-Related Potential Technique
Published on: September 7, 2022
08:43Combined Shuttle-Box Training with Electrophysiological Cortex Recording and Stimulation as a Tool to Study Perception and Learning
Published on: October 22, 2015