Context-FFT: A Context Feed Forward Transformer Network for EEG-based Speech Envelope Decoding
Ximin Chen, Yuting Ding, Nan Yan, Changsheng Chen, Fei Chen · 2024
Decoding speech envelope from electroencephalography (EEG) signals has been demonstrated to be useful for assessing speech intelligibility and boosting potential applications in neuroscience research as well as clinical diagnosis, which is also the focus of the ICASSP Auditory EEG 2023 Challenge. In order to further improve speech envelope decoding performance, this study proposes an end-to-end architecture based on multi-head attention mechanism called Context-FFT. Besides the transformer architecture, we also utilize a context layer to extract information and refine outputs based on given inputs. Notably, we decompose raw speech envelopes into envelopes in 12 frequency bands to model the relationship between 64-channel EEG signals and speech envelopes more precisely. Experiment results show that the proposed model achieves an average Pearson correlation value of 0.2148±0.1004 on held-out stories, outperforming the linear baseline by 51.65% and the VLAAI baseline by 23.17%, and 0.0701±0.0428 on held-out subjects. In terms of the final metric defined by the challenge, we obtain a final score of 0.1639, outperforming all submitted models of the ICASSP Auditory EEG 2023 Challenge. In the end, we explore the contributions of different brain regions to the process of continuous speech.