Validation of Speech Data for Training Automatic Speech Recognition Systems

Janez Križaj, Jerneja Žganec Gros, Simon Dobrišek · 2022 30th European Signal Processing Conference (EUSIPCO) · 2022

Recent automatic speech recognition systems are largely based on deep neural networks that need large amounts of labelled speech data to train. This can be a problem, especially for languages for which large speech databases are not available. To facilitate the construction of a speech database suitable for training automatic speech recognizers, we propose a tool that enables the validation of audio recordings from collected speech recordings. The developed tool allows the user to check the com-pliance with the predefined requirements regarding the correct audio format, the appropriate speech volume, the compatibility of the spoken text with the reference text and the suitability of the length of the non-spoken segments. The applicability of the developed tool is demonstrated by the creation of the Slovene speech corpus from audio recordings collected within the project Development of Slovene in a Digital Environment, although the tool is also suitable for all other languages supported by the used automatic speech recognizer.

Read the paper · More papers on PaperTik