Performance Comparison of Open Speech-To-Text Engines using Sentence Transformer Similarity Check with the Korean Language by Foreigners

Aria Bisma Wahyutama, Mintae Hwang · 2022

This paper contains the performance comparison of four Speech-to-Text (STT) engines which are Google STT, Naver Clova CSR, IBM Watson, and Microsoft Azure STT when transcribing foreigners speaking the Korean Language. The respondents are recording themselves speaking a predetermined sentence to be compiled together and then feeding it into the STT engine one by one to generate the transcribed text. The performance is evaluated using the Sentence Transformer Python framework that checks the similarity percentage between the original sentence to each of the transcribed texts and then finds the average result. The engine’s performance is categorized into four different categories which are sentence, nationality, age, and gender. The performance comparison results can be used to help determine the optimal STT engine for the Korean Language Spoken by Foreigner to develop STT-based or AI-based applications.

Read the paper · More papers on PaperTik