Managing Skew in Hadoop.

YongChul Kwon, Kai Ren, Magdalena Bałazińska, Bill Howe · 2013

Challenges in Big Data analytics stem not only from volume, but also variety: extreme diversity in both data types (e.g., text, images, and graphs) and in operations beyond relational algebra (e.g., machine learning, natural language processing, image processing, and graph analysis). As a result, any competitive Big Data system must support some form of parallel user-defined operations (UDOs) that can capture complex data processing tasks over complex data types without changing the core of the parallel data processing engine. Hadoop and other popular systems have been shown to provide a convenient programming model for implementing parallel UDOs, but the “black-box ” nature of UDOs complicates the automatic load balancing required to achieve parallel scalability. In this paper, we present an overview of some of our recent work that tackles the problem of load imbalance (a.k.a. skew) in parallel UDO evaluation. We first discuss the prevalence of skew in today’s applications and clusters. We then discuss our experience with static and dynamic methods for mitigating it.

Read the paper · More papers on PaperTik