Executing multiple group-by query in a MapReduce approach
Jie Pan, Frédéric Magoulès, Yann Le Biannic · 2010
Facing more and more generated information, data analysis software meets the challenge of processing large volume of data. The arrival of MapReduce provides a chance to utilize commodity hardware for processing large data set in parallel. In this paper, we focus on a special type of data analysis query, namely, multiple group-by query. We give an initial implementation of multiple group-by query based on MapReduce model. Considering the ignorable communication cost, we then propose an optimized version based on MapCombineReduce model, which addresses this issue. Our optimized version shows a better accelerating ability and a better scalability than the initial version.