Exploiting language-mismatched phoneme recognizers for unsupervised acoustic modeling

Siyuan Feng, Tan Lee, Haipeng Wang · 2016

This paper describes an investigation on acoustic modeling in the absence of transcribed training data. We propose to use language-mismatched phoneme recognizers to assist unsupervised segmentation and segment clustering of a new language. Using a language-mismatched recognizer, an input utterance is divided into many variable-length segments. Each segment is represented by a feature vector that is derived from the phoneme posterior probabilities. A spectral clustering algorithm is developed to group the segments into a prescribed number of clusters, which represent a set of basic speech units in the target language. By exploiting multiple recognizers for different languages, a wider phonetic space can be covered, leading to improved performance of segmentation and clustering. Experimental results on a multilingual speech database confirm the effectiveness of the proposed method.

Read the paper · More papers on PaperTik