Hadoop in Action: Building a Generic Log Analyzing System

Vinayak Bhosale, Anuja Thakar, Chinmaya Pandit, Anirudha Deshpande, Harmeet Kaur Khanuja · 2018

The massive amount of structured, semistructured and unstructured data can be attributed as Big Data. Traditional data processing applications and on-hand database system tools are impotent to evaluate and process such data. When the traditional systems were invented in the beginning, we never anticipated that we would have to deal with such monstrous amount of data, which would be difficult to process due to its characteristics viz. high volume, velocity, variety and veracity. The count of internet users has increased tremendously by virtue of which gargantuan amount of multifarious data gets generated. Moreover, the advent of IoT, has resulted in a boom in data generation. Thus, there was a pressing need to build a platform/framework which could process such huge, multifaceted data efficiently. That is where Hadoop came into picture. Hadoop facilitates scalable and distributed processing of Big Data in a proficient manner, saving a lot of time. Our work sheds some light on processing gigantic log data using Apache Hadoop open source framework. The data warehousing package built over Hadoop, Apache Hive, is used to summarize, analyze and query the different types of logs. Lastly, Apache Zeppelin, a powerful multipurpose notebook environment, is used for analysis, visualization and collaboration of log data in the form of bar graphs, pie charts, scatter charts, line graphs and area graphs.

Read the paper · More papers on PaperTik