Phoneme N-grams Based on a Polish Newspaper Corpus.
Bartosz Ziółko, Jakub Gałka, Mariusz Ziółko · International Conference on Artificial Intelligence · 2009
The phonetical statistics of Polish were collected from a newspaper corpus of around 110 000 000 words. The paper presents summarisation of the data which are phoneme ngrams and some phenomena in the statistics including a distribution of frequency of triphones occurring. Triphone statistics apply context-dependent speech units which have an important role in automatic speech recognition systems. The standard phonetic alphabet for Polish, SAMPA, and methods of providing phonetic transcriptions