Multi-Dialect Speech Recognition Using Transfer Learning and Transformer-Based Architectures: A Comprehensive Approach to Accurate and Efficient Dialect Identification
M. Malathi, S. Senthilkumar, CH Hussaian Basha, G. Sundaravadivel, M. Kavitha, Arunkumar P. · 2024
This paper presents a novel approach to multi-dialect speech recognition by leveraging transfer learning with transformer-based architectures, particularly Wav2Vec 2.0 and Whisper, to enhance ASR across dialectal variations. The model incorporates dialect-specific embeddings, allowing it to capture unique phonetic features and differentiate dialects efficiently. Data augmentation techniques like speed perturbation and pitch shifting are used to expand low-resource dialect data. Evaluation on multi-dialect datasets, including Common Voice and Arabic MGB-3, demonstrates significant improvements in Word Error Rate (WER) for dialects with limited resources, achieving effective zero-shot dialect recognition. This approach reduces dependence on extensive labeled datasets, offering practical applications in virtual assistants, transcription services, and call center automation in linguistically diverse regions.