Increasing Performance in Turkish by Finetuning of Multilingual Speech-to-Text Model

Öykü Berfin Mercan, Umut Özdil, Şükrü Ozan · 2022 30th Signal Processing and Communications Applications Conference (SIU) · 2022

This study was carried out with the aim of automatically translating phone calls between customers and customer representatives of a company. The dataset used in the study was created with audio files that were taken from open source platforms and reading of short texts in various contents by the company personnel. In addition to the labbeled data, approximately 28 thousand unlabeled data were labelled, and a total of 37534 audio data were prepared to be used in the training of the model that will translate from speech to text. The Wav2Vec2-XLSR-53 model which is a pre-trained model trained in 53 languages was fine-tuned with the our Turkish dataset. It has been obtained that it gives successful results in the speech to text performed on the data that is not used in model training and validation. The model was shared as open source on HugginFace to be used and tested for similar speech to text translation problems.

Read the paper · More papers on PaperTik