A Cleaning Algorithm for Noiseless Opinion Mining Corpus Construction
Otman Manad, Anna Pappa, Gilles Bernard · 2018
This paper presents DyCorC, an extractor and cleaner of web forums contents. Its main points are that the process is entirely automatic, language-independent and adaptable to all kinds of forum architectures. The corpus is built accordingly to user queries using expressions or item keywords as in research engines, and then DyCorC minimizes the boilerplate for further feature-based opinion mining and sentiment analysis, gathering comments and scorings. Such noiseless corpora are usually hand made with the help of crawlers and scrapers, with specific containers devised for each type of forum, entailing lots of work and skills. Our aim is to cut down this preprocessing stage. Our algorithm is compared to state of the art models (Apache Nutch, BootCat, JusText), with a gold standard corpus we released. DyCorC offers a better quality of noiseless content extraction. Its algorithm is based on DOM trees with string distances, seven of which have been compared on the reference corpus, and feature-distance has been chosen as the best fit.