Implementation and Performance Comparison of Partitioning Techniques in Apache Spark
J. Geetha, N. G. Harshit · 2019
Apache spark is one of the most demanded frameworks for High performance computing of Big Data. Data is growing day by day to such a large extent that the power of existing analytical tool is not sufficient. The degree of parallelism achieved directly impacts the performance of the framework. Parallelism can be achieved only when there is a proper division among the dataset. Partitioning helps in this division of work. Appropriate number of partitions and related data in partition will ensure the proper storage and faster access of data. Spark provides the inbuilt libraries and also flexibility to mould the available methods of partitioning according to convenience. In this paper we discus about the architecture of storage structure in spark, implementation of various partitioning techniques and performance comparison of those techniques.