MapReduce Tuning to Improve Distributed Machine Learning Performance
SungHwan Jeon, Haejin Chung, Wonseok Choi, Hee-Seong Shin, Jonghoon Chun, Jin Taek Kim, Yunmook Nah · 2018
In this paper, we show how MapReduce parameters affect distributed processing of machine learning programs, which are supported by machine learning libraries, such as Hadoop Mahout and Spark MLlib. We constructed virtualized cluster on top of Docker containers and measured distributed machine learning performance, while changing Hadoop parameters, such as number of replica, block size and memory buffer size.