Annotating Indirect Anaphora for Hindi: A Corpus Based Study

Pardeep Singh, Kamlesh Dutta · 2014

Natural language processing requires a lot of analysis and information regarding words and segment of sentence. Almost all NLP applications such as machine translation, information extraction, automatic summarization, question answering system, natural language generation, etc., require successful identification and resolution of anaphora. Information regarding word using POS tagger, parser and other tool can be gathered. Hindi is language of free word order as compare to English. This enforces additional constraints on different NLP task. In this working paper we present an analysis of Hindi genre. We used ten tags from literature. Out of ten tags seven are annotated using Botley's annotation scheme manually. We annotated 1540 demonstrative pronoun from twelve files of EMILEE corpus. Input file is EMILEE file and output is fully annotated unicode file.

Read the paper · More papers on PaperTik