A System for Unstructured Data Mining using Dynamic Ensemble Selection

Raquel Bezerra Calado, Leandro Sigfredo Rodriguez Torres, Alexandre Magno Andrade Maciel · 2020

Unstructured data represent as much as 90% of all business-relevant information. In Brazil, the practice of printing official journals dates back to the 19th century. Today more than 200 official journals in circulation, which together accumulate around 1.4 billion publications without textual standard. This work proposes the development of a system for unstructured data mining using a Dynamic Ensemble Selection. JudEasy implements, added in addition to classic text pre-processing methods, a set of twelve DES and a static method for creating categorized textual models for Brazilian of official journals. As results the DES-KL model obtained the highest accuracy rate of 96.81% and exceptional precision of 0.99.

Read the paper · More papers on PaperTik