A dynamic context shortening method for a minimum-context grapheme-to-phoneme data-driven transducer generator

Andrzej Pluciński · Journal of Quantitative Linguistics · 2006

We present an efficient way to learn automatically letter-to-phoneme mapping rules for Polish by using the concept of “dynamic context shortening method”.Attempts at reconstruction of transcription rules date back to 1987, when Sejnowski and Rosenberg applied a self-organizing neural network for “grapheme-to-phoneme” mapping. In the latter approaches decision tree based methods were applied. The trees for each letter were built starting from empty contexts. The left and right contexts were then alternately widened until the transcription ambiguity of the training data disappeared. We started in our approach from the symmetrical context wide enough to ensure unambiguous transcription in every context surroundings. Then both contexts, i.e., left and right, were shortened alternately until the ambiguity appeared. In all the cases where ambiguous transcriptions occurred, the previous context forms were restored and were not shortened further. Therefore at every step the cause of ambiguity, namely too short left or right context, was clearly known and removed. On the basis of the results obtained, the transcription tables for each letter were constructed. A 350,000 character corpus of Polish text transcribed into phonemic form was prepared and different-length training samples were taken from it at random and analysed. The remaining parts were used for verification. It turned out that it was enough to prepare a 30,000 character training sample to learn Polish grapheme-to-phoneme minimum-context mapping.We describe also three original generalization methods which we call rules coring, indeterminacies absorption and a guessing method. The last was invented to be applied in case of a limited acceptance of a context. The methods applied together allow the removal of about 70% of errors.

Read the paper · More papers on PaperTik