Application of artificial BEE colony algorithm using Hadoop

Nupur Bansal, Sanjay Kumar, Ashish Kumar Tripathi · International Conference on Computing for Sustainable Global Development · 2016

Big Data is the latest buzzword in IT industry especially for the last couple of years. Initially this word was used by organizations which had to deal with speedy growth of data such as web data, data which comes from scientific and business simulations or any other data resource. The Google File System and MapReduce Architecture have been the result of the pressure to handle continuously growing amount of data on the web (Ghemawat, Dean, & Sanjay, 2008). Evolutionary Computing (EC) is a stream of Computer Science that adapts the theory given as Darwinian evolution which is used to optimize many computing problems (Eiben & Smith, 2003). The concept of EC revolves around various natural evolution processes such as competition, reproduction, preying and random variation. Since evolution itself is an optimization process, thus the process of EC is applicable to a huge range of problems. One evolutionary algorithm is the Artificial Bee Colony (ABC) which is simple and effective, but this algorithm may take long hours or days to optimize difficult deceptive and/or expensive objective functions. ABC can be expressed naturally in Google's MapReduce framework (Ghemawat, Dean, & Sanjay, 2008) to develop a simple and robust implementation which is parallel as well that includes communication, fault tolerance and load balancing. Since this implementation is flexible it is easy to do some modifications to it, which can improve optimization of objective functions and also improve parallel performance. In this paper, we propose and realize an approach to implement Artificial Bee Colony Algorithm in a parallel manner and use it to compute the effort of COCOMO Model. The population of the bee here are in the form of big data are stored in Hadoop Distributed File System (HDFS) on Hadoop to have better fault tolerance and to store large amount of data in distributed manner.

Read the paper · More papers on PaperTik