Summarization using mapreduce framework based big data and hybrid algorithm (HMM and DBSCAN)

Krushnadeo T. Belerao, S. B. Chaudhari · 2017 IEEE International Conference on Power, Control, Signals and Instrumentation Engineering (ICPCSI) · 2017

With the adventure of the internet there is vast increase in the storage of information. Almost all the information exists in digital form which reduces lots of the paper work and increases ease of storage. Searching relevant information in collection of documents is a tedious task. The solution comes in picture for this problem is automatic text summarization. In this paper, abstract summary generation of large amount of documents for big data is proposed which will consider user input as topic. The proposed system is designed using DBSCAN algorithm which works with MapReduce framework for clustering and Hidden Markov Model for summarization. The summarization process is performed in three main stages and provides a modular implementation of large numbre of documents for summarization. The proposed method follows preprocessing step in which documents are scanned with similarity and various machine learning technique are applied. The result of applying clustering enhances the summarizer system to collect exact words rather than copying redundant words. Topic based abstract summarization from big data is challenging task particularly when there are multiple documents with same or different content. Hadoop with its programming techniques can provide better ways of generating summary and it also enhances the complexity of summarization process using distributed computing.

Read the paper · More papers on PaperTik