Grapheme to Phoneme Conversion of Norwegian using Hidden Markov Models

Terje Solsvik Kristensen, Markus Sauer Nilssen · 2020

In this paper, the applicability of Hidden Markov Models to the grapheme-to-phoneme (GTP) problem of Norwegian is explored. The grapheme-to-phoneme problem, is part of the problem of converting sequences of graphemes to sequences of phonemes. This is an important issue in both text-to-speech and speech recognition systems. With the assistance of established toolkits like CMU-Cambridge Language Modeling Toolkit and Hidden Markov Model Toolkit (HTK), an approach based on Hidden Markov Models (HMM) is presented and implemented. By such an approach every phoneme is modeled by a HMM that generates the graphemes. This approach has previously been tested on English data. By using HMM for Norwegian, a phoneme accuracy of 94%, and a word accuracy of 68% is obtained. This is slightly better than similar results obtained for English.

Read the paper · More papers on PaperTik