Performance Evaluation of Big Data Technology on Designing Big Network Traffic Data Analysis System
Nattawat Khamphakdee, Nunnapus Benjamas, Saiyan Saiyod · 2016
Network and computer systems administrators are facing a serious problem of the big network traffic data analysis. It became difficult work of administrators to extract and analysis the abnormal and normal patterns from large amounts of the network traffic data. Currently, traditional relational database management systems (RDBMS) are unsuitable to store a large amount of data because they are designed for storing and processing the structured data. Hive is a data warehouse tool built on top of Hadoop for storing, processing, querying, and analysis the large amount of data. Hive stores the data in a table similar the relational database management system. In this paper, we propose a Hadoop-based traffic querying and analyzing system that handles the TCP, ICMP, and UDP analysis of the big network traffic data. The system consists of six modules: Data Collection Module, Transferring and Storing Information Module, Convertor Module, Data Mining Process Module, DM2SC Module, and Report Module. We also performed complex search queries and compared the query response times of MySQL against Hive in Hadoop environment. As the result, in some scenario, MySQL outperform a cluster of four Hive nodes on querying the ICMP protocol information, nevertheless, MySQL database that stored more than the network traffic data about 45 million records cannot be query the TCP protocol information. Moreover, we observed that the average query response times of Hive in Hadoop cluster that reduce continuously be scale up node into the cluster.