Comparative Analysis of Hadoop Tools and Spark Technology
Aniket Wakde, Purvesh Shende, Sudarshan Waydande, Shravani Uttarwar, Ganesh Deshmukh · 2018
Data continuously goes on increasing day by day due to rapidly increasing population, use of sensors, use of social media, use of Internet of things etc. Data generated from these various sources may be small or it may be large, also it can be structured or unstructured. We can easily process and analyze small amount of data in a single system. But nowadays data generating through various sources is big in size and mostly in unstructured form. So, in order to analyze such big amount of unstructured data we need some mechanism which can handle this amount of data and process it. Hadoop and Spark are the technologies which can handle any kind of data and can process it. For storing, processing and analyzing the data, different technologies and tools are used under hadoop and spark ecosystem. In this paper, we specifically focus on use of Hadoop tools and technologies like Map-Reduce, Apache flume, Apache Pig and Apache Spark technology and their comparative analysis.