Evaluating the Performance of Turkish Automatic Speech Recognition Using the Generative AI-Based Whisper Model
Tunahan Gökçimen, Bihter Daş, Resul Daş · 2024
Automatic speech recognition (ASR) for the Turkish language faces significant challenges due to its agglutinative structure and diverse phonetic variations. In this study, we evaluate the performance of OpenAI's Whisper models of varying sizes-tiny, base, small, and medium-on Turkish ASR tasks. The experiments were designed to measure each model's training loss, validation loss, word error rate (WER), and training time. Our findings indicate that larger models, specifically the whisper-medium, achieve superior performance with the lowest validation loss and WER, albeit at the cost of longer training times. The results demonstrate a substantial improvement in Turkish ASR capabilities compared to existing models, filling a significant gap in the literature. This study not only advances the state of ASR for the Turkish language but also provides valuable insights into the trade-offs between model complexity and performance, guiding future research and applications in the field.