Optimized Multiple Platforms for Big Data Analysis
Bao Rong Chang, Hsiu Fen Tsai, Yo-Ai Wang · 2016
The objective of this study is to optimize a multiple big data processing platform with high performance and high availability. The optimization to the integration of Apache Hive, Cloudera Impala and BDAS Spark SQL enables the platform to support SQL queries in a big data environment, automatically select the best performing big data warehouse platform for computing, and receive the same result far more rapidly from the high-performance cache system. The proposed approach significantly improves overall performance, especially in terms of the application of multiple repeated SQL commands in multi-user mode, thus dramatically reducing the query/response time in such scenarios.