Phonetic Modeling in the Philips Chinese Continuous-Speech Recognition System

Frank Torsten Bernd Seide, Nick J.-C. Wang · 1998

We have extended the Philips large-vocabulary continuousspeech recognition system towards Chinese. On the way from our existing Western-language technology to Mandarin, the rst step was to build a suitable phonetic model. This paper describes the development of our phonetic model (excluding tones) for Mandarin Chinese. We will present asystematic comparison of three forms of sub-syllabic units for Chinese, phonemes, initials / nals, and a non-tonal form of preme/toneme models, aswell as wholesyllable models for reference. We include experiments on bottom-up and decision-tree based top-down state clustering and modelling of cross-syllable contexts. All forms of sub-syllabic units are represented in the Philips Mandarin phone set \\SAMPA-C. " SAMPA-C is based on the European SAMPA standard and introduced in this paper. Our studies show that traditional half-syllable approaches slightly outperform Western-style triphones. Modelling of right-context dependency gives greater improvement than left-context dependency, and cross-syllable modelling yields a 4-5 % performance gain. In a free syllable decoding task, we achieve 39 % syllable error rate for telephone speech and 24 % for microphone dictations. 1.

Read the paper · More papers on PaperTik