Big Data Analytics: An Approach using Hadoop Distributed File System

P. Beaulah Soundarabai, S Aravindh, J Thriveni · 2014

Today's world is driven by Growth and Innovation for a better future. All of which are based on analysis and harnessing of tons of data, typically known as Big Data. The tasks involved for achieving results at such a scale can be challenging and painfully slow. This paper works towards an approach for effectively solving a large and computationally intensive problem by leveraging the capabilities of Hadoop and Hbase. Here we demonstrate how to reduce and distribute the large problem across multiple nodes as small chunks and later aggregate the results obtained for the various chunks to arrive at the desired result. This method of distributing and aggregating is called map-reduce, wherein dividing and distributing the problem to different nodes is called Map task and aggregation of results is called Reduce task. As part of implementation we will be analyzing population census data. In information technology, big data is a collection of data over a period of time, and can exist in both structured and unstructured form. The data could be anything from pictures, videos and posts from a social media site to climate information collected by weather sensors. This data is so large and complex that it is a challenge to extract, clean, transform, store, analyze and harness useful information. This exercise falls beyond the scope of a RDBMS or traditional data processing applications, simply due to the existence of tons of bytes of data. Yet, it is imperative that we work with such Big Data to derive correlations and other information to aide development, growth, break-through and innovation in various fields. Almost every domain is reliant on information that can be harnessed from Big Data to help move towards progress. (1 to 5) . Big Data can be typically characterized by the three V's -  Volume - Exponentially growing data.  Variety - Data comes in all shapes, sizes and forms.  Velocity - Rate of data creation and the rate at which data is analyzed and harnessed for useful information. The paper proposes an implementation of Big Data analytics using Hadoop and Hbase. It is important to note that the tests conducted are infrastructure dependent and may vary based on the underlying hardware and network. Motivation: Big data analysis with the help of famous tools like informatics, vertical etc., help us to get the details of the entire data, get an insight and to make decisions based on the results derived. Most of the tools are not freely available and they require training and maintenance from the experts on the particular tools. We wanted to bring out the benefits of Hadoop freamework which is an open source tool and show how it reduces the implementation time and the cost with respect to the centralized system. Contribution: In this paper, we have proposed a Big data analysis with HDFS Framework of Hadoop that works towards providing better performance in terms of time, cost and user friendliness. The main objective of our proposed model is to apply business logic on a huge volume of data. We have run eight different queries on 2GB data size of audio conferencing call files and estimated the runtime of the queries. We have applied filtering of the selected attributes and on the queries were made to run on the projected structure. Organization: The rest of the paper is organized as follows: Section II shows the related literature of other researchers; Background of Hadoop system is present in Section III. Problem Definition and the Methodology are available in Section IV; Section V describes the implementation and the results; Section VI details the Performance Evaluation; Conclusion of the work is presented in Section VII.

Read the paper · More papers on PaperTik