Survey of Caching Mechanism in Hadoop
Jalpa Shah, Hinal Somani · International journal of advance research and innovative ideas in education · 2016
Hadoop is open source project by apache which can be used as stand alone, single node and multi node cluster. Hadoop is develop in version 1x to 2x with map-reduce v2 and YARN by the Apache devlopers.Hadoop is mainly consist of HDFS and Map-Reduce, where HDFS is as storage for Hadoop and Map-Reduce is programming model for data file processing. Hadoop has a centralize cache management which is useful for repetitively accessed files. The distributed cache copies files to every node, then map or reduce reads the files from the local file system and make the data any time accessible any time for use as it is distributed to multiple location in the DataNode blocks of fixed sized.