Enabling open domain information exploitation for disaster management

Bharathi Ramudu, Malay Kumar Nema, R. P. Ram Kumar · 2017

This work proposes an efficient method for the data pre-processing step of the information cycle. The method considers the open domain news feeds as data source. We try to handle the issue of open domain news data becoming unmanageably huge by designing a hybrid schema. The schema consists of two fold processing. One sifting and other filtering, both at logical levels. The data feeds are subjected to selection of the relevant data using DOM parser as first step. Subsequently the captured data size is reduced by filtering the irrelevant text from the HTML page using HTML parser. After the proposed pre-processing we get highly relevant data which in turn can enhances the effectiveness of the various post processing algorithms. Disaster management is taken as an application scenario. While designing the pre-processing method, the operational constraints of disaster management were considered, which we try to address. The results of the experiments shows significant data size reduction while keeping the relevant details intact.

Read the paper · More papers on PaperTik