Comparative Analysis of Speech Transcription Technologies for the Digitalization of Technical Support Services

Zoya V. Arkhipova, Valery Staver · System Analysis & Mathematical Modeling · 2025

This article is dedicated to the application of neural network technologies to enhance the efficiency and quality of technical support services. The use of speech transcription technologies is becoming increasingly relevant due to the rising demands for high-quality information processing in various fields. The study examines the main approaches to speech transcription, including classical methods, deep learning-based solutions, hybrid approaches, as well as commercial and open-source tools. The research aims to conduct a comparative analysis of modern transcription systems to select and subsequently implement the most effective solution in the technical support service of the franchising company “Laboratory S,” as employees of the company currently manually record conversations with clients after each interaction. For the purposes of the study, both commercial solutions and open-source tools were utilized. Commercial systems (Google Speech-to-Text, Yandex SpeechKit, Amazon Transcribe, Azure Speech-to-Text) were applied directly via the official platforms of the respective services. Open-source solutions (Kaldi, DeepSpeech, OpenAI Whisper) were deployed in the Google Colab environment. The transcription results obtained from both commercial and open-source tools were then compared within Google Colab, where Python and libraries such as NumPy and scikit-learn were used to calculate metrics and assess transcription quality. The effectiveness of these systems was evaluated using the WER (Word Error Rate), MER (Match Error Rate), WIP (Word Information Preserved), and WIL (Word Information Loss) metrics. Two datasets were used for analysis: the first consisted of recordings made under ideal conditions, and the second included recordings reflecting the real working scenarios of “Laboratory S.” The results of the study revealed the strengths and weaknesses of various speech transcription technologies and determined their applicability in the real working environment of a technical support service.

Read the paper · More papers on PaperTik