JOUM: An Indexing Methodology for Improving Join in Hive Star schema
Hussien Sh, Abdel Azez, Mohamed H. Khafagy, Fatma A. Omara · International Journal of Scientific and Engineering Research · 2015
New applications such as machine learning, web searches, recommendation engines, and social networks generate enormous amounts of logs, email, and other technical structured/unstructured information streams. Such applications need fast processing of data which is involved in today’s business processes analysis. These applications might contain several thousand tables with over hundreds of terabytes of data which are used heavily for both reporting and decision-making which often don't need update or deletion operations [1]. Map Reduce is a programming model for large-scale distributed data processing with simple and elegant concepts which are used to build blocks for other parallel programming tools. In the same time, it is considered extensible for different applications providing advantages of Concurrency/Parallelism, tolerating failures and hiding any complexity from the user. So, Map Reduce has become the important standard for large-scale data processing in many enterprises. Also, it is used for developing new solutions on massive datasets such as relational data analytics, web analytics, machine learning, real-time analytics and data mining [2, 3]. Hadoop is considered a framework based on Map Reduce programming model for large-scale distributed data processing. According to Hadoop, the applications run on large clusters. These clusters are built from a variety of homogeneous hardware. They provide the applications both reliability and data mobility. Therefore, Hadoop implements Map Reduce concepts Where the application is divided into small tasks, every task could run or re-executed on any cluster's node. Also, Hadoop uses a distributed file system that stores data on the compute nodes as a tree of distributed blocks, and provides data