Automatically learning the units of speech by non-negative matrix factorisation
Veronique Stouten, Kris Demuynck, Hugo Van hamme · 2007
We present an unsupervised technique to discover the (word-sized) speech units in which a corpus of utterances can be de-composed. First, a fixed-length high-dimensional vector rep-resentation of the utterances is obtained. Then, the resulting matrix is decomposed in terms of additive units by applying the non-negative matrix factorisation algorithm. On a small vocab-ulary task, the obtained basis vectors each represent one of the uttered words. We also investigate the amount of speech data that is needed to obtain a correct set of basis vectors. By de-creasing the number of occurrences of the words in the corpus, an indication of the learning rate of the system is obtained. Index Terms: matrix factorisation, word segmentation, phone lattices, language acquisition.