Unsupervised acoustic corpora building based on variable confidence measure thresholding

Tomáš Koctúr, Ján Staš, Jozef Juhár · 2016

In order to achieve better automatic speech recognition results, more complex acoustic models are required. Acoustic models are usually trained on hundreds or thousands of hours of transcribed speech data. Obtaining training data is very time consuming process and it is a huge problem for under-resourced languages. Unsupervised acoustic model training procedures have been used for years now but they have a lot of limitations and quality issues. The main problem of unsupervised acoustic model training is that the data selected for training process contain a lot of errors and that is the reason why resulting acoustic models are not very precise. The goal of our research is to build and tune an unsupervised acoustic model training system, which will generate corpora in very high quality (approx. 1% word error rate) in order to achieve very good acoustic model. In this article, in order to obtain more acoustic data in very good quality, filtration technique based on variable confidence measure thresholding is proposed.

Read the paper · More papers on PaperTik