Turkish Speech Recognition for Voice Search Applications

Hilal Tekgöz, Muhammed Murat Özbek, Tolga Büyüktanır, Harun Uz · 2024

Nowadays, with the development of technology, studies in the field of natural language processing are gaining importance. Speech recognition systems, one of the natural language processing applications, are one of the active areas of study. Voice search engines have the ability for users to issue commands via voice. Voice recognition models are used in voice search engines. In this study, fine-tuning was carried out on pre-trained Whisper-Large, Seamleess and Wav2vec 2.0 Base models, which were developed specifically for Turkish speech recognition and have multi-language support, in order to provide voice search features to search engines. The resulting models were evaluated with word error rate (WER) and character error rate (CER) evaluation metrics. According to the findings, the highest performance rate in the models used to convert voice to text in our voice search system was obtained with the SeamlessM4T model. Although Wav2vec 2.0 Base was pre-trained with the original dataset, it did not produce results as successful as the SeamlessM4T and Whisper-Large models.

Read the paper · More papers on PaperTik