End-To-End Tibetan Speech Recognition Method Integrating Feature Extraction and Data Augmentation
Weiyuan Zhang, Dondrub Lhakpa, Droma Dechen, Xiahui Yi · 2025
The rapid advancement of artificial intelligence has made speech recognition a critical research direction in information technology. However, Tibetan speech recognition faces challenges such as limited research capacity and insufficient resources, leading to low accuracy and slow inference. This study addressed these issues by adopting an end-to-end model framework to improve inference efficiency while integrating data augmentation and feature extraction. Statistical analysis of annotated Tibetan letters was performed to construct an table for acoustic model training. Fbank method extracted effective audio features, converted into digital inputs to accelerate training in the end-to-end architecture. Cepstral mean and variance normalization was applied to audio data to standardize features. Five data augmentation were implemented to enhance signal robustness, complemented by a bidirectional Transformer decoder and CTC algorithm for optimized decoding. Experimental results show that in the end-to-end Tibetan speech recognition model, the effective combination of these methods significantly improves performance. Experiments on 34-hour Ü-Tsang dialect speech data set achieved 4.93% character error rate.