BenCo: First Step Towards Coreference Resolution in Bengali
Mojammel Hossain, MD Samiul Islam, Jashim Uddin Ahmed, Md Rahat Kader Khan · 2023
Coreference resolution is well-studied in NLP; however, Bengali coreference resolution research has not been as well investigated as it has been for English and other rich languages. Bengali has a richer morphology than English despite having fewer resources. This research introduce a tiny new dataset of coreference annotations over Bengali texts from four domains in this article called BenCo. In this dataset, there are 48,610 tokens and 5488 mention annotations structured into 470 mention clusters. The article outline the procedure used to generate this dataset and use it to build an end-to-end neural network-based system. This study will help to clarify how Bengali coreference phenomena differ depending on the domain and inspires others to produce more materials in the language. Also, insufficient information transfer in the zero-shot multilingual situations, which may indicate the need for language-specific resources for this job.