Map reduction framework for parallel data mining: Multicore to distributed network systems

Rahul Ramakrishna, M V Bhaskara Rao · 2011

In this multi core era, there is a huge influx of symmetric multi-process computers based on shared memory architecture and high end server platforms. It appears no adequate framework exists to manifest the complete potential of the hardware. In this paper a parallel programming framework is demonstrated applicable to different algorithms in a distinctive way from the conventional single algorithm speedup at a particular point of time. The framework fosters application dependent speedup over uniprocessor applications for a given workload and even on small Ethernet/IP based networks. Functional programming paradigm has the ability to implicitly parallelize program to multicore computers and scaled in distributed networks using a message queues. Also the map reduction framework is based on functional programming paradigm, where the programs can be written in summation form, specifying a map function which generates intermediate key value pairs and a reduce function merging the key value pairs. With this method a substantial increase in speed efficiency is obtained. However, the framework by itself will not substantially increase the speed of execution, as other parameters like chunking of data affect the performance metrics. Graphical methods are used and explained in order to show the optimum amount of chunking to be used for execution.

Read the paper · More papers on PaperTik