Lightweight Transformer based end-to-end speech recognition with patch adaptive local dense synthesizer attention
Peiyuan Tang, Penghua Li, Shengwei Liu, HaoXiang Xu · 2024
This paper proposes a novel Transformer model, called PALDSA-Transformer, for automatic speech recognition (ASR) tasks. In the model, the patch adaptive local dense synthesizer attention (PALDSA) is designed to capture the local context correlations and reduce attention complexity. In this paper, we study how to reduce the transformer model architecture complexity with a limited computing budget, leading to a more efficient architecture design.We introduce the Convolution Subsampling Module to the encoder and replace redundant pre-layer Normalization layers with a scaled post-layer Normalization in the Convolution Subsampling Module and Feed Forward Module. Compared with other variants of transformer-based methods, the experimental results show that the proposed method achieves a lower word error rate (WER) of 3.51%/8.82% for test-clean/test-other without a language model for only 8.6M parameters on the Librispeech dataset.