Parallelization with standard MLA, a standard ML API for hadoop and comparison with existing large-scale solutions

Ngọc Tuấn Nguyễn · NORA - Norwegian Open Research Archives · 2015

In recent years, the world has witnessed an exponential growth and availability of data.The term "big data" has become one of the hottest topics which attracts the investment of a lot of people from research to business, from small to large organizations.People have data and want to get insights from them.This leads to the demand of a parallel processing models that are scalable and stable.There are several choices such as Hadoop and its variants and Message Passing Interface.Standard ML is a functional programming language which is mainly used in teaching and research.However, there is not much support for this language, especially in parallel model.Therefore, in this thesis, we develop a Standard ML API for Hadoop called MLoop to provide SML developers a framework to program with MapReduce paradigm in Hadoop.This library is an extension of Hadoop Pipes to support SML instead of C++.The thesis also conducts experiments to evaluate and compare proposed library with other notable large-scale parallel solutions.The results show that MLoop achieves better performance compared with Hadoop Streaming which is an extension provided by Hadoop to support programming languages other than Java.Although its performance is really not as good as the native Hadoop, MLoop often gets at least 80% the performance of the native Hadoop.In some cases (the summation problem for example), when strengths of SML are utilized, MLoop even outperforms the native Hadoop.Besides that, MLoop also inherits characteristics of Hadoop such as scalability and fault tolerance.However, the current implementation of MLoop suffers from several shortcomings such as it does not support job chaining and global counter.Finally, the thesis also provides several useful guide-lines to make it easier to choose the suitable solution for the actual large-scale problems.

Read the paper · More papers on PaperTik