Comparison of Spark Resource Managers and Distributed File Systems
Soulmaz Salehian, Yonghong Yan · 2016
The rapid growth in volume, velocity, and variety of data produced by applications of scientific computing, commercial workloads and cloud has led to Big Data. Traditional solutions of data storage, management and processing cannot meet demands of this distributed data, so new execution models, data models and software systems have been developed to address the challenges of storing data in heterogeneous form, e.g. HDFS, NoSQL database, and for processing data in parallel and distributed fashion, e.g. MapReduce, Hadoop and Spark frameworks. This work comparatively studies Apache Spark distributed data processing framework. Our study first discusses the resource management subsystems of Spark, and then reviews several of the distributed data storage options available to Spark.