Development of the Co-reference Resolution Tagged Data set in Assamese @ A Semi-Automated Approach

Mridusmita Das, Apurbalal Senapati · 2023

Co-reference resolution (CR) is an essential task in several Natural Language Processing (NLP) applications. It implies finding all the linguistic expressions (known as mentions) in a given text and the references of those mentioned. Consider the following text: Ishwar Chandra Bandyopadhyay was a great man in India in the nineteenth century. He received the honorable title Vidyasagar. In this text, each italics phrase is a mention and all are co-references. That means both expressions He and Vidyasagar refer to the same discourse entity Ishwar Chandra Bandyopadhyay. It is used in various NLP applications like Text Summarization, Information Retrieval (IR), Question Answering systems, Machine Translation (MT), etc. To develop a co-reference system using any algorithm must be needed the data. For a resource-scarce language like Assamese, there is no such data set or supportive tools to prepare the data. So, the data preparation is the primary task of the research in co-reference resolution. This paper attempt a semi-automated approach to develop a tagged data set for the co-reference resolution. This resource will boost the researchers in NLP research in Assamese languages.

Read the paper · More papers on PaperTik