Research on Serialization Storage Strategy Based on Spark Cluster

Fangfang Yang, Yuchong Xia · 2017

Spark is a kind of big data processing platform based on memory computing.The Spark default serialization strategy has low utilization of cache which has greatly influenced the efficiency of Spark task execution.For solving this problem of low computational efficiency caused by insufficient memory, this paper proposes an optimized serialized storage strategy, which combining with the running cot of RDD, RDD execution time and count of Action.Experimental results show that the proposed strategy can improve the computational efficiency under the limited task resources.

Read the paper · More papers on PaperTik