SparkOT: Diagnosing Operation Level Inefficiency in Spark
Honggang Zhou, Yunchun Li, Jie Jia, Weichen Qi, Hailong Yang · 2018
Spark is a popular big data framework aiming to speed up computation by caching data in memory. Spark splits the execution of an application into multiple operations which form a directed acyclic graph. Operations that only depend on the data partitions from their parents are grouped into a task. Traditional performance analysis of big data framework mostly studies straggler task and proposes mitigation technique which ignores the deeper reasoning about the efficiency of the internal operations. In addition, task level analysis cannot correlate ill-behaved task to the specific source code. In this work, we present a tool that dynamically instruments Spark to discover inefficient operations without changing the source code of both the framework and the application. In addition, we propose a new score function for each operation within a stage, which synthetically considers the impact of pipeline duration, peer operation and operation duration. Moreover, we also propose a characterization method that performs clustering analysis on the sampled JVM features in order to better understand the cause of the inefficient operations and provide useful guidance for performance optimization.