Smart-MLlib: A High-Performance Machine-Learning Library
David Siegal, Jia Guo, Gagan Agrawal · 2016
As the popularity of big data analytics has continued to grow, so has the need for accessible and scalable machine-learning implementations. In recent years, Apache Spark's machine-learning library, MLlib, has been used to fulfill this need. Though Spark outperforms Hadoop, it is not clear if it is the best performing underlying middleware to support machine learning implementations. Building on a C++ and MPI based middleware system Smart, we present a machine-learning library prototype (Smart-MLlib). Like MLlib, Smart MLlib allows machine learning implementations to be invoked from a Scala program, and with a very similar API. To test our library's performance, we built four machine-learning applications that are also provided in Spark's MLlib: k-means clustering, linear regression, Gaussian mixture models, and support vector machines. On average, we outperformed Spark's MLlib by over 800%. Our library also scaled better than Spark's MLlib for every application tested. Thus, the new machine-learning library enables higher performance than Spark's MLlib without sacrificing the easy-to-use API.