Optimizing Deep Neural Networks for EEG-Based Speech Recognition: A Multimodal Approach to Assistive Communication
A deep learning-based multimodal framework that combines EEG and audio signals to decode spoken phrases. By integrating neural activity with speech audio, the system improves recognition for individuals with neurological impairments, such as Parkinson's disease, ALS, or cerebral palsy, and enhances robustness in noisy environments. The framework allows more inclusive and accurate speech recognition, enabling better access to voice-0driven technologies and assistive communication tools.
Traditional automatic speech recognition (ASR) systems struggle with irregular speech patterns and noisy environments, resulting in high error rates and limited accessibility for individuals with speech impairments. Many users with neurological conditions are unable to benefit from voice -controlled applications due to these limitations. There is a need for a multimodal approach that incorporates EEG-derived neural features alongside audio to improve recognition accuracy and inclusivity.
This University at Buffalo technology employs a multimodal deep learning architecture that fuses processed EEG and audio features for speech decoding. The audio encoder is a GRU-based model that processes MFCC features to produce a latent representation of speech, while the EEG component includes three variations to evaluate different representations. The Time-Distributed CNN extracts frequency-based spatial features from EEG windows, the ConvLSTM2D combines convolutional layers with LSTM to capture spatial-temporal dependencies, and the GRU-based EEG encoder models temporal patterns in statistical EEG features. Latent representations from both audio and EEG are then combined through a late fusion strategy, processed via fully connected dense layers with ReLU activation and dropout, and passed to a softmax classification head. By integrating spectral, temporal and spatial features from both modalities, this framework enables robust classification of spoken phrases and allows assessment of the relative importance of different EEG features in decoding speech.
Header image is purely illustrative. Source: psdesign1, stock.adobe.com
- Enhances speech recognition for individuals with neurological impairments.
- Robust performance in noisy environments.
- Leverages complementary neural and audio information for improved accuracy.
- Flexible encoder architrectures allow optimization for different EEG representations.
- Supports inclusive, assistive communication applications.
- Assistive communication devices for individuals with speech impairments.
- Voice-controlled interfaces for users with neurological conditions.
- Speech recognition in noisy or real-world environments where traditional ASR fails.
- Research into neural decoding of speech and multimodal brain-computer interfaces.
Provisional patent application 63/984,828 filed February 17, 2026
Validated with experimental EEG and audio datasets, prototype neural network models demonstrated (TRL 4-5).
Available for licensing or collaboration.
Patent Information:
| App Type |
Country |
Serial No. |
Patent No. |
Patent Status |
File Date |
Issued Date |
Expire Date |
|