BigData and MapReduce with Hadoop
Roman Trobec, Aleksandra Rashkovska, P Mežnar, Andrej Lipej · 2012
The paper describes the application of the MapReduce paradigm for processing and analyzing large amounts of data, coming from a computer simulation for a specific scientific problem. The Apache Hadoop open source distribution was installed on a cluster built of six computing nodes, each with four cores. The implemented MapReduce job pipeline is described and the essential Java code segments are presented. The experimental measurements of the employed MapReduce tasks execution times result in a speedup of 20, which indicates that a high level of parallelism is achieved.