Evaluating the impact of data placement to spark and SciDB with an Earth Science use case

Khoa D. Doan, Amidu Oloso, Kwo‐Sen Kuo, Thomas L. Clune, Hongfeng Yu, Brian R. Nelson, Jian Zhang · 2016

We investigate the impact of data placement on two Big Data technologies, Spark and SciDB, with a use case from Earth Science where data arrays are multidimensional. Simultaneously, this investigation provides an opportunity to evaluate the performance of the technologies involved. Two datastores, HDFS and Cassandra, are used with Spark for our comparison. It is found that Spark with Cassandra performs better than with HDFS, but SciDB performs better yet than Spark with either datastore. The investigation also underscores the value of having data aligned for the most common analysis scenarios in advance on a shared nothing architecture. Otherwise, repartitioning needs to be carried out on the fly, degrading overall performance.

Read the paper · More papers on PaperTik