{FMEM}: A Fine-grained Memory Estimator for MapReduce Jobs
Lijie Xu, Jie Liu, Jun Fang Wei · International Conference on Autonomic Computing · 2013
MapReduce is designed as a simple and scalable framework for big data processing. Due to the lack of resource usage models, its implementation Hadoop hands over resource planning and optimizing works to users. But users also find difficulty in specifying right resource-related, especially memory-related, configurations without good knowledge of job’s memory usage. Modeling memory usage is challenging because there are many influencing factors such as framework’s dataflow, user-defined programs, large space of configurations and memory management mechanism of JVM. In order to help both users and the framework to analyze, predict and optimize memory usage, we propose a Fine-grained Memory Estimator for MapReduce jobs called FMEM. FMEM contains a dataflow estimator which can predict the data volume flowing among map/reduce tasks. Based on dataflow and rules of memory utilization learnt from a lot of jobs, FMEM uses a rules-statistics method to estimate fine-grained memory usage in each generation of task’s JVM. Representative benchmarks show that FMEM can predict diverse jobs’ memory usage within 20% relative error. Furthermore, FMEM will be promoted to find optimum dataflow and memory related configurations.