Related Experiment Video
Updated: Jan 14, 2026

08:15
Capturing Dynamic Finger Gesturing with High-resolution Surface Electromyography and Computer Vision
Published on: March 28, 2025
1.2K
A Dual-Modal Silent Speech Interface via Surface Electromyography (sEMG) and Vibration Sensing
Guang-Yang Gou1, Yu-Sen Guo2, Zi-Xuan Song1
1State Key Laboratory of Transducer Technology, Aerospace Information Research Institute (AIR), Chinese Academy of Sciences, Beijing 100190, China.
ACS Applied Materials & Interfaces
|October 23, 2025
Summary
This study introduces a dual-modal silent speech recognition system using hydrogel surface electromyography (sEMG) and polyvinylidene fluoride (PVDF) vibration sensors. The innovative system achieves robust, privacy-preserving human-machine communication.
Area of Science:
- Biomedical Engineering
- Materials Science
- Signal Processing
Background:
- Conventional acoustic speech recognition struggles with noisy environments and vocal impairments.
- Nonacoustic physiological signals offer an alternative for speech recognition.
- Need for robust and privacy-preserving human-machine interfaces.
Purpose of the Study:
- To develop a dual-modal silent speech recognition system using nonacoustic throat signals.
- To enable robust speech acquisition resistant to environmental noise and vocal issues.
- To create a scalable framework for privacy-preserving human-machine communication.
Main Methods:
- Fabrication of hydrogel-based surface electromyography (sEMG) electrodes using poly(vinyl alcohol) (PVA)/poly(acrylic acid) (PAA).
- Integration of a polyvinylidene fluoride (PVDF) piezoelectric vibration sensor.
- Co-integration of sEMG and PVDF sensors onto a polyimide (PI) substrate for colocalized sensing.
- Development of a dual-branch silent interaction network (DSI-Net) for multimodal signal fusion and decoding.
Main Results:
- Hydrogel sEMG electrodes achieved low interfacial impedance (270 Ω at 1 kHz) for high-fidelity biopotential recording.
- PVDF sensor demonstrated high sensitivity (365 mV/Pa at 300 Hz) and a wide frequency response (50 Hz-1 kHz) with SNR > 48 dB.
- The DSI-Net significantly improved silent speech recognition accuracy by fusing multimodal inputs.
Conclusions:
- The dual-modal system effectively captures neuromuscular and vibrational throat signals for silent speech recognition.
- This technology offers a scalable and robust solution for privacy-preserving human-machine communication.
- Potential applications include assistive technologies and resilient voice interfaces for challenging environments.

