Pronunciation Lexicon Development for Under-Resourced Languages Using Automatically Derived Subword Units: A Case Study on Scottish Gaelic
Marzieh Razavi, Ramya Rasipuram, Mathew Magimai.-Doss · Infoscience (Ecole Polytechnique Fédérale de Lausanne) · 2015
Developing a phonetic lexicon for a language requires linguistic knowledge as well as human effort, which may not be available, particularly for under-resourced languages.To avoid the need for the linguistic knowledge, acoustic information can be used to automatically obtain the subword units and the associated pronunciations.Towards that, the present paper investigates the potential of a recently proposed hidden Markov model formalism for automatic derivation of subword units and lexicon development on a truly under-resourced and endangered language, more precisely Scottish Gaelic.Our studies show that the formalism can not only be useful in developing a lexicon that helps in building better automatic speech recognition systems, but can also be extended to find the relationship between the derived subword units and the existing knowledge about phonetic units from resource-rich languages, more precisely multilingual phones.Thus, the formalism paves a path for systematically combining acoustic and linguistic knowledge from multiple languages with the limited acoustic and linguistic knowledge of the under-resourced language in order to develop phone-like automatic subword unit based lexical resources.