A Novel Architecture to Integrate Multi-Source Data into Distributed Environment using Big Data Infrastructure
Sidra Zulfiqar, Amen Faridoon, Muhammad Imran · 2019
The amount of data has been increasing over the last few years due to the emergence of various end-user applications. These applications utilize cloud computing infrastructure in the data centers. Apart from the increasing volume of data, there are other factors such as variety, velocity, and veracity of the data which result in the problem of big data. Traditional database management systems are not efficient to handle big data. The use of big data platform is necessary to resolve the big data problem. Hadoop is one of the platforms which resolves the problem of big data. Hadoop uses distributed storage. HBase is one of the big data tools for storing big data in Hadoop. HBase is a column-oriented, distributed and high fault-tolerant database. It can store billions of rows at a time. However, there are some issues in HBase. When the data comes from multiple sources, it is stored in multiple tables in HBase. As a result, the performance of HBase degrades when there is a need to perform join operations. In this research, we propose a data model which stores data of multiple sources into a single HBase table. There is no issue of join performance of HBase in the proposed technique as there is no need to perform join because the data is integrated into a single HBase table. The proposed model has a unique row key and multiple column families to integrate data from multiple sources in single table. We evaluated the proposed technique using a real testbed by considering a dataset of two publishers. We compare the performance of the proposed technique by storing data in to multiple tables in Hive. Results show improved query performance of the proposed technique as compared to the traditional approach of using join operations in multiple tables in Hive.