A profiling tool for apache hive run-time query
Divya Kamath, Praveen Srinivas, Ashika Gopal, B. V. Lanchana, V. Suma · 2017
Apache Hive is a tool used conventionally for data warehousing and analysis. Although it is widely used, there is very less research on performance comparison and analysis. One of the reasons is, the techniques applied to supervise execution cannot be implemented to intermediate MapReduce code developed from Hive query. Since the MapReduce code is hidden from the developer, the only opportunity to view the actual execution is through run time logs. It is essential to develop an automatic tool for generating and extracting reports from runtime logs to understand the query execution behavior. A tool is designed to build an execution profile of individual hive queries for information extracted from hive and Hadoop logs. Information about MapReduce jobs, tasks and attempts belonging to a query are present in the profile. This data is saved in MongoDB as a JSON document which can be extracted as charts or table. Several experiments were conducted on queries to demonstrate that the profiling tool is capable of assisting developers for comparison of hive queries of different format and different parameters. Performance issues can be diagnosed by comparing the attempts/tasks within the same job.