SMR: Scalable MapReduce for Multicore Systems

Yu Zhang, Yufen Yu, Jiankang Chen · 2018

Although multicore chips have been widely used, it is still a challenge on how software fully utilizes multicore resources. MapReduce originally designed for clusters alleviates the burden of programmers by providing automatic parallelism. However, its implementations for multicore systems cannot scale up the performance by adding more CPU cores. Experimental results show that kernel spinlock on shared address space is the main reason for performance downgrade when core count exceeds a specific number, e.g., 8 or 16. To address the issue, this paper proposes a multithreaded model Sthread which provides isolated address spaces between threads to avoid contentions, and provides unbounded-channel abstraction for asynchronously passing unbounded data streams between threads. Based on Sthread, a scalable MapReduce library SMR is proposed, which breaks the phase barrier and improves performance by privatizing data buffers and adopting unbounded-channels between phases. Experimental results show that SMR achieves better scalability and performance than Phoenix. Specially, performance improvements range from 9× to 26.7× for hist, wc, pca on 32 cores.

Read the paper · More papers on PaperTik