SparkArray: An Array-Based Scientific Data Management System Built on Apache Spark

Wenjuan Wang, Taoying Liu, Dixin Tang, Hong Liu, Wei Li, Rubao Lee · 2016

With the highly demanded requirements for manipulating large scientific datasets, scientists are in need of flexible cluster-level software to execute fast scientific data analysis. In this paper, we discuss whether the Apache Spark framework is suitable for scientific data management. We present our system SparkArray, which extends Spark with a multidimensional array data model and a set of common used array operations (e.g., filter, subarray, smooth and join). We present analysis and performance evaluation results on different implementation methods for executing array operations on clusters. We also compared the performance of SparkArray to SciDB, a recent scientific database management system, using the workloads of the Standard Science DBMS Benchmark (SS-DB). The results show that SparkArray is feasible alternative solution for large-scale scientific data management, especially when scientists require fast data loading or one-time data analysis on large scientific datasets.

Read the paper · More papers on PaperTik