Survey on high performance analytics of bigdata with apache spark

Ramkrushna C. Maheshwar, Dasari Sai Naga Haritha · 2016

This paper lays attention upon the advantages of Apache Spark over Hadoop MapReduce and analysis of real time data using time-series analysis. As Hadoop MapReduce is a widely used and famous execution engine for working with the storage and analysis of large datasets. In MapReduce, the data is read from the disk and the result is written to the Hadoop Distributed File System (HDFS) after a particular iteration and then the data is read from the HDFS for the next iteration. This whole process consumes a lot of disk space and time as well. The users had been objecting the problem of high latency and fault tolerance of the entire system. To overcome the issues and disadvantages of MapReduce, Apache Spark was developed. Apache Spark is an open-source project that ensures lower latency queries, iterative computations and real time processing on similar data. This paper also focuses on time-series analysis in Hadoop and Spark environment which processes and does analysis of real-time data and generates a pattern out of it to get a clearer glimpse of the statistics and characteristics of data thus making Spark even more efficient over MapReduce.

Read the paper · More papers on PaperTik