Machine Learning approach for resolving anaphors in the Punjabi Language

Kawaljit Kaur, Vishal Goyal, Kamlesh Dutta · 2021

The paper presents an attempt to resolve anaphors in the Punjabi language using Machine Learning strategy. The Punjabi language has limited resources and no corpus is available which can be used for the task, so we prepared our own. Punjabi Shallow Parser has been used as a tool to provide information from the text such as part of speech (POS) tags, number, person, gender of NPs and chunking information. Animacy, NER, Pronoun Type features and linkages of anaphors to associated entities are manually annotated on the data. The features retrieved from the annotated corpus were used to train different classifiers, which were then evaluated using Precision, Recall, and F-Score. Analysis of the results shows that ensemble classifiers have performed better than other classifiers. We also tested the contribution of different features in resolving anaphors by training classifiers with different subsets of features. The accuracy of the pre-processing tool(Punjabi Shallow Parser), and the quality of the annotation task completed are also the factors that affect the system's performance. Because this is the first attempt to resolve anaphora in Punjabi, the results could be improved by using a larger corpus and adopting a rule-based strategy in conjunction with a machine learning approach, which is the future scope of the current work.

Read the paper · More papers on PaperTik