Document Summarization in Malayalam with sentence framing
Kavya Kishore, Greeshma N. Gopal, P H Neethu · 2016
Document Summarization is a technique of conveying important information in a given document. It is one of the most important chores of Natural Language Processing as the summary produced is helpful for information retrieval systems, question answering systems, medical domain and news domain etc. Most of the summarization works in Indian languages are of extractive nature and not much work is oriented towards the abstractive summarization approaches in Indian languages as they need more linguistic processing. As Indian languages belong to several language families and since they are morphologically rich and agglutinative in nature, a lot of challenges are faced while doing abstractive summarization because it requires natural language generation techniques. The proposed work is an approach for abstractive text summarization that will accept single document as input in Malayalam, processes the input by building a suitable semantic representation and then use sentence framing techniques to generate the final summary. The entire framework is composed of eight modules that mainly deals with constructing a suitable semantic representation called Karaka tree and a sentence framing module to generate the natural summary.