Dependency Treebank of Urdu and its Evaluation

Riyaz Ahmad Bhat, Dipti Misra Sharma · 2012

In this paper we describe a currently un-derway treebanking effort for Urdu-a South Asian language. The treebank is built from a newspaper corpus and uses a Karaka based grammatical framework inspired by Paninian grammatical theory. Thus far 3366 sen-tences (0.1M words) have been annotated with the linguistic information at morpho-syntactic (morphological, part-of-speech and chunk in-formation) and syntactico-semantic (depen-dency) levels. This work also aims to evalu-ate the correctness or reliability of this man-ual annotated dependency treebank. Evalua-tion is done by measuring the inter-annotator agreement on a manually annotated data set of 196 sentences (5600 words) annotated by two annotators. We present the qualitative analy-sis of the agreement statistics and identify the possible reasons for the disagreement between the annotators. We also show the syntactic annotation of some constructions specific to Urdu like Ezafe and discuss the problem of word segmentation (tokenization). 1

Read the paper · More papers on PaperTik