Unsupervised submodular subset selection for speech data

Kai Wei, Yuzong Liu, Katrin Kirchhoff, Jeff Bilmes · 2014

We conduct a comparative study on selecting subsets of acoustic data for training phone recognizers. The data selection problem is approached as a constrained submodular optimization problem. Previous applications of this approach required transcriptions or acoustic models trained in a supervised way. In this paper we develop and evaluate a novel and entirely unsupervised approach, and apply it to TIMIT data. Results show that our method consistently outperforms a number of baseline methods while being computationally very efficient and requiring no labeling.

Read the paper · More papers on PaperTik