Related Experiment Video
Updated: Mar 27, 2026

fMRI Mapping of Brain Activity Associated with the Vocal Production of Consonant and Dissonant Intervals
Published on: May 23, 2017
Soft Multiaxial Strain Mapping Interface with AI-Driven Decoding for Silent Speech in Noise
Sunguk Hong1, Junyoung Yoo2, Sung-Min Park1,2,3,4,5,6
1Department of Mechanical Engineering, Pohang University of Science and Technology (POSTECH), Pohang 37673, South Korea.
Abstract:
Silent speech interfaces (SSIs) offer a viable alternative to traditional microphones in capturing clear audio in noisy environments. We propose a reconceptualized SSI that reproduces voice by monitoring continuous multiaxial strain maps induced by throat muscle movements. The system integrates a computer vision-based optical strain (CVOS) sensor with deep learning-based voice reconstruction, enabling clear alphabetic communication under extreme noise conditions. The CVOS sensor-comprising a soft silicone substrate with micromarkers and a tiny camera-achieves high-sensitivity marker detection and captures complex strain patterns with higher scalability and reliability compared to conventional wearable sensors. The inference pipeline of the CVOS-based SSI incorporates physics-based automated baseline calibration and content-adaptive temporal attention, enabling robust analysis of the captured strain patterns. Based on the inference results, a personalized text-to-speech model subsequently reconstructs the speaker's voice. These algorithmic features ensure robustness under dynamic conditions by employing real-time adaptive signal processing that compensates for inter- and intrasubject anatomical variability. Alphabet-based communication is achieved through the synergy between optimized algorithms and interface design. The performance of the CVOS-based SSI was validated in real-world noisy scenarios, confirming its practical applicability.

