How to Process Bangla Linguistic-Data through NLP Pipeline

Mohammad Mamun Or Rashid · Dhaka University Journal of Linguistics · 2021

Raw texts are the most used applied form of human language, especially stored in electronic and digital media. In this research, we would like to propose a Bangla language processing with annotation guidelines from raw data to value-added data. The main objective of this paper is defining all input-output specification which is pipelined as processing phases. The major processing phases are set on the text understanding part. Test understanding part requires various types of tagging and labeling tasks like PoS tagging, parsing, NER tagging, etc. Moreover, another two steps need to be achieved to comply with the pipeline, that is coreference resolution and word sense disambiguation which define the semantic states of any linguistic input.

Read the paper · More papers on PaperTik