Stimulus Reconstruction Based Auditory Attention Detection Using EEG in Multi-Speaker Environments Without Access to Clean Sources

Kai Yang, Xueying Luan, Gaoyan Zhang · 2022

Auditory attention detection (AAD) based on electroencephalography (EEG) can be utilized to help hearing-impaired people improve speech perception abilities in multi-speaker environments. Previously, most studies relied on clean speech sources to perform EEG-based AAD, which limits the practical applications in life scenarios with mixed speech and noise. In this paper, we first proposed using a dual-path recurrent neural network (DPRNN) to separate individual speech signals in a multi-speaker scenario and then applied an EEG-based speech reconstruction model by a temporal convolutional network (TCN) to perform AAD. By comparing the correlation of the reconstructed speech envelopes from EEG with the actual ones, the attended side is detected by a higher correlation than the unattended side. Experiments on an EEG dataset with no directional information of speech show that the AAD accuracies of our model are improved for both short and long detection windows (2-s and 5-s) compared with previous leading reconstruction-based AAD methods. Analysis of different frequency bands shows that the contribution of the $\beta$ band EEG signal to AAD is greatest in both the 2-s and 5-s models. Visualization of the weights suggests that the temporal, frontal, and parietal lobes of the brain contribute most to the model. These findings are consistent with previous studies on the neural mechanism of AAD, which demonstrates the effectiveness and interpretability of the proposed model.

Read the paper · More papers on PaperTik