Towards Conceptual MapReduce Algorithm for Big Data Platform

Seungdae Sohn, Jin-Hong Kim · 2015

MapReduce Comes from its simplicity to preparing the input data, the programmer needs only to implement the mapper, the reducer, and optionally, the combiner and the partitioner. All other aspects of execution are handled transparently by the execution framework on clusters ranging from a single node to a few thousand nodes, over datasets ranging from gigabytes to petabytes. However, this also means that any conceivable algorithm that a programmer wishes to develop must be expressed in terms of a small number of rigidly defined components that must fit together in very specific ways. It may not appear obvious how a multitude of algorithms can be recast into this programming model. The purpose of this paper is to provide, a guide to MapReduce algorithm design. This paper presents the notion of design pattern of MapReduce, which instantiate arrangements of components and specific techniques designed to handle frequently encountered situations across a variety of domains.

Read the paper · More papers on PaperTik