Towards Creating a Speech Corpus for Polish Using Public Domain Audiobooks

Artur Zygadło, Artur Janicki · 2019

The goal of automatic speech recognition (ASR) is to convert spoken language into text. To deliver high-quality speech recognition, large training data sets, called speech corpora, are required. Preparation of such resources is time-consuming and usually involves many people. In this paper, an attempt to generate a Polish speech corpus automatically, based on existing public domain audiobook recordings and book texts, is presented. The described system for automatic generation of speech corpora was designed and then implemented with the Kaldi toolkit and additional scripts. Several experiments were conducted to prove the usefulness of the obtained corpus in speech recognition and to emphasize the importance of the amount of training data.

Read the paper · More papers on PaperTik