Computational linguistics resources for Indo-Iranian languages
Shafqat Mumtaz Virk · Gothenburg University Publications Electronic Archive (Gothenburg University) · 2013
Can computers process human languages? During the last fifty years, two main approaches have been used to find an answer to this question: data-driven (i.e. statistics based) and knowledge-driven (i.e. grammar based). The former relies on the availability of a vast amount of electronic linguistic data and the processing capabilities of modern-age computers, while the latter builds on grammatical rules and classical linguistic theories of language. In this thesis, we use mainly the second approach and elucidate the de-velopment of computational (”resource”) grammars for six Indo-Iranian lan-guages: Urdu, Hindi, Punjabi, Persian, Sindhi, and Nepali. We explore different lexical and syntactical aspects of these languages and build their resource grammars using the Grammatical Framework (GF) – a type theo-retical grammar formalism tool. We also provide computational evidence of the similarities/differences between Hindi and Urdu, and report a mechanical development of a Hindi resource grammar starting from an Urdu resource grammar. We use a func-tor style implementation that makes it possible to share the commonalities between the two languages. Our analysis shows that this sharing is possible upto 94 % at the syntax level, whereas at the lexical level Hindi and Urdu