A Multi-Head Attention-Based Spatiotemporal Convolutional Model for Decoding Auditory EEG Evoked by Mandarin Tones
Qi Zhang, Qingxin Meng, Xue Zhang, Lan Tian · 2025
Auditory electroencephalogram (aEEG) can objectively assess a subject's auditory perception ability, and the Mandarin four tone is a crucial syllable feature in speech interaction. This paper proposes a deep learning method based on Multi-Head Attention Spatiotemporal Convolution (MATCNet) model which can high-performance decode and classify 4-tone aEEG signals evoked by Mandarin syllable. Under a specific EEG paradigm, when listening to the same syllable with different tones, aEEG-evoked signals were acquired for each subject. For the raw aEEG signals, the preprocessing steps including filtering, downsampling, segmentation, and denoising, were operated. The presented MATCNet model is based on a sliding-window attention-based temporal convolutional network architecture which consists of multiple modules, including a convolutional network (CV), a multi-head attention mechanism (MAT), and spatiotemporal convolution (TC). On the dataset included 1,800 aEEG samples, the model super-parameters were trained and optimized. The experimental results showed that the MATCNet method can achieve higher accuracy comparing to other two deep learning models, and the average accuracy of 4-tone aEEG decoding is up to 91.11%.