NoSQL vs relational database: A comparative study about the generation of the most frequent N-grams

Jardel Ribeiro, Jonas Henrique, Rodrigo Ribeiro, Rosalvo Ferreira de Oliveira Neto · 2017

This study intends to help data mining developers to get better performance when obtaining the most frequent N-grams in Text Mining projects. The process of building new variables is one of the oldest and still challenging problems in Data Mining projects. The most frequent N-grams are commonly used as input variables in Text Mining projects. The N-grams represent the occurrence of N items in sequence in a given text. The items can be letters or words. This paper presents a performance comparison between the two main approaches of data storage, relational and NoSQL databases in the task of obtaining the most frequent N-grams. Validation of the study was executed using a database from a known benchmark from an international competition organized by PAN@CLEF 2013. The one-tailed paired t-test showed that NoSQL approach is statistically superior to the relational approach with a confidence level of 95%.

Read the paper · More papers on PaperTik