Resilient distributed computing platforms for big data analysis using Spark and Hadoop

Bao Rong Chang, Hsiu Fen Tsai, Yo-Ai Wang, Chien‐Feng Huang · 2016

This paper introduces the integration of three platforms using Apache Hive, Cloudera Impala and BDAS Spark SQL which enables to support SQL-like queries in big data environment. In order to fast respond to user's query for big data processing, the optimized system can automatically select the appropriate platform to best perform a query. In addition, the rapid data retrieval from the in-memory cache or in-disk cache has achieved for the repeated SQL command. The proposed approach improves the efficiency of data retrieval significantly.

Read the paper · More papers on PaperTik