Adaptive Mask Based Attention Mechanism for Mandarin Speech Recognition
Penghua Li, Jiawei Cheng, Yujun Rong, Ziheng Huang, Xiao Meng Xie · 2022 34th Chinese Control and Decision Conference (CCDC) · 2022
This paper proposes a novel Transformer model, called ALDSA-Transformer, for extracting local and global information in Mandarin automatic speech recognition (ASR) tasks. In such a model, the adaptive local dense synthesizer attention (ALDSA) is designed to capture the local context correlations. The span of each attention head in ALDSA is parameterized by introducing the adaptive mask function. After training, the adaptive mask function with fixed attention span is applied to the attention matrix for masking the corresponding values. Besides, the ALDSA is integrated with self-attention (SA) to build the ALDSA encoder block. Furthermore, with the stacked encoder blocks, the ALDSA-Transformer encoder is constructed to extract the acoustic features in Mandarin ASR task. Compared with other variants of transformer-based methods, the experimental results show that the proposed method achieves a lower character error rate (CER) of 5.65% with parameter size 23.0M on AISHELL-1 dataset.