Enhancing Speech-to-Text Transcription Accuracy for the Bahraini Dialect
Abdulla Almahmood, Hesham Al-Ammal, Fatema A. Albalooshi · 2024
This study investigates means of enhancing the accuracy of Automated Speech Recognition (ASR) systems for the Bahraini dialect, a variant that has received minimal attention in natural language processing research. This is achieved through increasing the accuracy of the OpenAI Whisper model's transcription for the Bahraini dialect. Two tailored audio datasets were created: one with a local Bahraini audio which was manually transcribed, and another that included a wider range of audio sources such TV broadcasts, podcasts, and parliament sessions. To optimize the Whisper model, several methods were used, such as sequential training on both datasets, data augmentation, and hyperparameter optimization with Optuna. The Word Error Rate (WER) significantly decreased from 181.608% to 13.54% in the results, demonstrating the efficacy of the fine-tuning techniques in raising the model's accuracy above its original benchmarks.