Optimizing Join in HIVE Star Schema Using Key/Facts Indexing
Hussien SH. Abdel Azez, Mohamed H. Khafagy, Fatma A. Omara · IETE Technical Review · 2017
These days Big Data represents complex and an important issue for the extraction/retrieval of information due to the fact that its analysis requires massive computation power. In addition, database star schema can be considered as one of the complicated data models due to the use of joining queries heavily for information extraction and reports generation, which demands scanning for a large amount of data (tera, peta, zeta bytes, etc.). On the other hand, HIVE is considered one of the essential and efficient Big Data SQL-based tools built on the top of Hadoop as a translator from SQL queries into Map/Reduce tasks. In addition, using data indexing techniques with join queries could improve /speed up HIVE join query tasks execution especially in a star schema. According to the work in this paper, Key/Facts indexing methodology was introduced to materialize the star schema and inject a simple index for data. Based on this, Key/Facts indexing methodology SQL queries’ execution time in HIVE improved without changing HIVE framework. TPC-H benchmark was used in order to estimate the performance of Key/Facts methodology. Experimental results prove that Key/Facts methodology out-performs traditional HIVE join execution time. Also, Key/Facts performance is improved by increasing the data size. Generally, Key/Facts can be considered one of the suitable methodologies for Big Data analysis.