Sensitivity analysis of latency to data size in Spark environment

Abdelaziz El Yazidi, Moulay Lahcen Hasnaoui, Mohamed Saad Azizi · 2020

The explosion of quantity of data coming from the Internet of Things (IoT) and the generation of big data in social networks and E-commerce websites, becomes a difficult task for researchers because of the 5 V characterizing this type of data (volume, variety, variability, veracity and velocity). Many platforms have been proposed to deal with big data in the WSN domain. This article focuses on studying the performance of the Big data environment platform, namely Spark. We start by describing the Big Data ecosystem, then we will attack the components of the Spark platform, namely the Hadoop Distributed file system for storage, and MapReduce with YARN (Yet Another Resource Negotiator) for parallel processing, thus, we will describe the collection of libraries that uses resilient distributed data to overcome computational complexity. The last part of this article is devoted to an experiment which we will apply a MapReduce job which will calculate the number of replication of a word in files of different sizes with the Spark environment and we will analyze the results in speed of execution of complex calculations (in our case browsing a file of more than 6 million lines) by job MapReduce, we obtained results which favor the veracity of big data according to the size of data received by the cluster.

Read the paper · More papers on PaperTik