Related Experiment Video
Updated: Sep 5, 2025

05:48
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
1.6K
CNN Architectures and Feature Extraction Methods for EEG Imaginary Speech Recognition
Ana-Luiza Rusnac1, Ovidiu Grigore1
1Department of Applied Electronics and Information Engineering, Faculty of Electronics, Telecommunications and Information Technology, Polytechnic University of Bucharest, 060042 Bucharest, Romania.
Sensors (Basel, Switzerland)
|July 9, 2022
Summary
This study developed a low-cost imaginary speech recognition system using frequency domain covariance and convolutional neural networks (CNNs). The system achieved 37% accuracy, demonstrating effective communication for individuals with neural dysfunctions.
Area of Science:
- Biomedical Engineering
- Computer Science
- Neuroscience
Background:
- Neural dysfunctions significantly impair speech, impacting daily communication.
- Developing accessible assistive technologies is crucial for individuals with communication challenges.
- Imaginary speech recognition offers a potential solution for non-vocal communication.
Purpose of the Study:
- To design and optimize an intelligent imaginary speech recognition system for low-cost applications.
- To identify optimal parameters for feature extraction and neural network architecture.
- To evaluate system performance on a limited-resource platform.
Main Methods:
- Utilized the Kara One database for phoneme and word recordings.
- Employed frequency domain covariance for feature extraction, outperforming time-domain methods.
- Investigated various window lengths (0.25s, 0.5s, 1s) and convolutional neural network (CNN) architectures.
- Tested on eight subjects, aiming for a subject-independent system.
Main Results:
- Frequency domain covariance with a 0.25s window achieved the best performance.
- A CNN with two convolutional layers (64/128 filters) and a dense layer (64 neurons) yielded optimal results.
- Achieved up to 37% accuracy for 11 phonemes and words.
- The system demonstrated a low-cost profile with a 1.8 ms running time on an AMD Ryzen 7 4800HS CPU.
Conclusions:
- Short-term signal analysis is vital for imaginary speech recognition.
- Complex CNN architectures do not guarantee superior performance.
- The developed system is suitable for low-cost, resource-limited applications, offering a viable communication aid.
Keywords:
Kara One databaseconvolutional neural networkelectroencephalographyimaginary speechsignal processing
