Performance analysis of MySQL partition, hive partition-bucketing and Apache Pig
A. Sunny Kumar · 2016
Streaming data analysis has attracted attention In various applications like financial records, data analysis, etc. Such type of applications require continuous storage of large amount of data in data warehouse while simultaneously providing quick response time for the queries against the data that is stored in the system. The duration of fetching data varies depending on type of data required from the system. This paper presents the performance estimates in terms of MySQL Partition, Hive partition-bucketing and Apache Pig framework. In this paper, big data eco systems and comparative performance analysis of frequently used data retrieval techniques such as MySQL, Hive and Pig are described. From the work presented in the paper, it is concluded that the execution time for extracting data becomes very large with growth in data size, particularly in case of MySQL. As compared to MySQL, Hive and Pig takes less time and give better results.