Robust Imagined Speech Production from Electrocorticography with Adaptive Frequency Enhancement
Xu Xu, Chong Fu, Junxin Chen, Gwanggil Jeon, David Camacho · 2024
Imagined speech production with electrocorticography (ECoG) plays a crucial role in brain-computer interface system. A challenging issue is the great variation underlying the frequency bands of the ECoG signals’ encode information, which makes current methods difficult to generate imagined speech with stable quality among different persons. To this end, we propose a robust model to generate high-quality imagined speech from ECoG. A frequency enhancement branch is first designed to adaptively modulate the frequency information, whose product is fed into the following multi-scale channel attention module for robust feature extraction and fusion. By incorporating both the mel-spectrum and audio as training constraints, a multi-constraint decoder branch is finally constructed for imagined speech production. The performance of our model is evaluated on a high-quality dateset, i.e, Single Word Production Dutch-iBIDS. It yields Pearson correlation scores that are all above 0.8, and the standard deviationsare are all below 0.2 in different volunteers. Experimental results demonstrate that our model is effective and robust for ECoG based imagined speech production, and has advantages over peer methods.