Natural Language Processing based Rule Based Discourse Analysis of Marathi Text
Kalpana B. Khandale, C. Namrata Mahender · 2020 International Conference on Electronics and Sustainable Communication Systems (ICESC) · 2020
Indian languages are verb final and morphologically rich. For the present work Marathi language is explored which belongs to Indo Aryan family. The data is collected from the Marathi Balbharati book from standard 1st to 4th. This paper presents challenges of discourse analysis means two or more sentences are linked together which is required for many natural language based applications. Rule based structure is framed to find out the discourse from the dataset. These rules are framed on the basis of the Marathi grammar and world knowledge of the linguistic. With the help of POS tagger identify the tag of each word in the discourse is identified. The discourse anaphora is resolved in various foreign as well as Indian languages but Marathi is one of the languages where the discourse is not resolved uptil now. Discourse means that the some relative sentences or the sub sentences are connected with each-other. An important step in understanding the connected discourse is skill to link the meaning of the successive sentence together. Total 515 discourse sentences are considered here to resolving the anaphora from which 177 are the simple sentence and 338 are the discourse one.