Processing time of TFIDF and Naive Bayes on Spark 2.0, Hadoop 2.6 and Hadoop 2.7: Which Tool Is More Efficient?

Erin Gilheany · TRAP@NCI (National College of Ireland) · 2016

There has been a large emphasis placed on the performance variations which occur when comparing Hadoop and Spark. This research paper will dive into the details of this comparison using TD-IDF and Naive Bayes algorithms on both applications to demonstrate the total processing time differences. It has been noted in literature that Spark goes a long way towards dealing with the limitations of Hadoop, in particular those issues which frequently arise in the application of iterative machine learning algorithms due to the slow processing of inputs/outputs to disc. This paper explores the difference from a text categorisation stand point. On a single computer there is a strong distinction in computing times of TFIDF on Spark Versus Hadoop, with Spark completing the application in a fraction of the time. Naive Bayes shows a contrasting picture with Spark's processing speeds on average twice as big as that of Hadoop. Given the additional costs of RAM on Spark, in this instance Hadoop would appear to be the better choice.

Read the paper · More papers on PaperTik