Evaluating the impact of data placement to spark and SciDB with an Earth Science use case
Khoa D. Doan, Amidu Oloso, Kwo‐Sen Kuo, Thomas L. Clune, Hongfeng Yu, Brian R. Nelson, Jian Zhang · 2016
We investigate the impact of data placement on two Big Data technologies, Spark and SciDB, with a use case from Earth Science where data arrays are multidimensional. Simultaneously, this investigation provides an opportunity to evaluate the performance of the technologies involved. Two datastores, HDFS and Cassandra, are used with Spark for our comparison. It is found that Spark with Cassandra performs better than with HDFS, but SciDB performs better yet than Spark with either datastore. The investigation also underscores the value of having data aligned for the most common analysis scenarios in advance on a shared nothing architecture. Otherwise, repartitioning needs to be carried out on the fly, degrading overall performance.