Developing a finite-state morphological anlayzer for Urdu and Hindi

Tina Bögel, Miriam Butt, Annette Hautli-Janisz, Sebastian Sulger · Finite-State Methods and Natural Language Processing · 2007

We introduce and discuss a number of issues that arise in the process of building a finite-state morphological analyzer for Urdu, in particular issues with potential ambiguity and non-concatenative morphology. Our approach allows for an underlyingly similar treatment of both Urdu and Hindi via a cascade of finite-state transducers that transliterates the very different scripts into a common ASCII transcription system. As this transliteration system is based on the XFST tools that the Urdu/Hindi common morphological analyzer is also implemented in, no compatibility problems arise.

Read the paper · More papers on PaperTik