Web data harvesting for speech understanding grammar induction
Ioannis Klasinas, Alexandros Potamianos, Elias Iosif, Spiros Georgiladakis, Gianluca Mameli · 2013
The development of a grammar for a spoken dialogue system can be greatly accelerated by using a corpus describing the application. However the development of such a corpus is a slow and expensive process. This paper proposes unsupervised methods for finding relevant corpora in the Web and mining the most informative parts. We show that by utilizing perplexity we are able to increase the in-domainess (precision) of the mined corpora, while by utilizing the rank of the web search engine we can increase the generalizability (recall). The results show that using only unsupervised and language independent methods we can compete with corpora created manually with expert knowledge.