Parallel bulk insertion for large-scale analytics applications

Antonio Barbuzzi, Pietro Michiardi, Ernst W. Biersack, Gennaro Boggia · 2010

Modern data analytics applications, e.g. Internet-scale indexing, system trace analysis, recommender engines to name a few, operate on massive amounts of data and call for a parallel approach to data processing. In this work, we focus on the popular MapReduce framework to carry out such tasks and identify bulk data insert operations as a critical preliminary step to achieve reduced processing times, especially when new data is generated and processed at regular time intervals.

Read the paper · More papers on PaperTik