A Data Aware Caching for Large scale Data Applications Using The Map-Reduce
Rupali V. Pashte · 2014
The Big-data refers to the huge scale distributed data processing applications that operate on unusually large amounts of data. Google’s MapReduce and Apache’s MapReduce, its open-source implementation, are the defacto software systems for Large Scale data applications. Study of the MapReduce framework is that the framework generates a large amount of intermediate data. Such existing information is thrown away after the tasks finish, because MapReduce is not able to utilize them. In this paper, we propose, a data-aware cache framework for large data applications. In this paper, tasks submit their intermediate results to the cache manager. A job queries the cache manager before executing the actual computing work. A novel cache description scheme and a cache request and reply protocols are designed. We implement Data aware caching by extending Hadoop.