Lite ASR Transformer: A Light Weight Transformer Architecture For Automatic Speech Recognition
N J Metilda Sagaya Mary, Srinivasan Umesh · 2024
Transformers are popular sequence-to-sequence models but have large number of parameters and high compute requirements. As an initiative to reduce the energy demand by Transformer models and to get better Transformer models for edge devices, we propose a light weight Transformer in this paper. We attempt to reduce the compute and carbon footprint of the original Transformer architecture by incorporating architectural modifications. The proposed modifications reduce the Transformer parameters by 42.7 % relative to the original Transformer of same depth and width. Our automatic speech recognition experiments on LibriSpeech, SPGISpeech and GigaSpeech datasets show that the proposed light weight Transformer has negligible ASR performance degradation. The compute requirements also reduce by 23% relative to the original Transformer of same depth and width. We also show that the proposed Lite ASR Transformer has acceptable convergence and also the latency is 20% lesser relative to the original Transformer of same depth and width.