Robust Imagined Speech Production Using AI-Generated Content Network for Patients With Language Impairments
Xu Xu, Chong Fu · IEEE Transactions on Consumer Electronics · 2024
Imagined speech production is critical for brain-computer interface systems. It is able to provide the communication ability for patients with language impairments. Nowadays, many studies have developed algorithms for this purpose. However, they often struggle with an issue that is the noise in physiological signals from various devices. This makes current methods have low robustness when performing speech generation. To this end, we propose a framework to robustly producing imagined speech for patients with language impairments. It is based on artificial intelligent-generated content (AIGC). In specific, this work develops a diffusion model to remove noise from brain signals, which are then converted and normalized into log-mel spectrograms. Next, the product feeds into a transformer-based model to extract local and global features. Finally, these enhanced features are inputted into full connected layers to generate the imagined speech. The effectiveness of our model is validated using the high-quality, i.e., single-word-production-Dutch-iBIDS dataset. It achieves Pearson correlation scores above 0.8 and standard deviations below 0.2 across different volunteers. Experimental results well demonstrate that our model is robust and has advantages over peer methods.