A Novel Technique for Parallelization of Genetic Algorithm using Hadoop

L. M. R. J. Lobo · 2013

Document categorization is used in education, government sectors, art, industry etc. Categorizing a document to enable immediate finding of it in the future motivated the concept of Classification involving Document categorization. Manual document classification involves a lot of effort and is time consuming. The basic idea implemented in this paper speeds up processing and reduces manual intervention, by atomizing this categorization. This idea is an edge over the existing classification systems. The implementation of the system basically involves getting into to parallelize Genetic Algorithm (GA) thus improving the processing speed. The use of Hadoop MapReduce and HDFS (Hadoop Distributed File System) framework helps to store big data and speeds up the calculations involved in the computation of genetic algorithm. The motivation of this work has reason from mapreduce fare well in terms of scalability, fault tolerance, and ease-of-use. This is adjoined by hadoop being an open-source and hadoop being written in Java. Parallel genetic algorithm may be implemented by many methods and produced better results than sequentially run GA in terms of space and time. The approach expressed in this paper enables us to develop Parallel Genetic Algorithm (PGA). Cloud computing framework could be used. This has been recognized as one of the latest computing paradigm where applications, data and IT services are provided over the Internet (1). A plugin to data mining packages OlexGA is used for developing classification. Hadoop technology along with its components can be used to store big data and process computations at a faster rate. Hadoop provides various technologies like MapReduce framework, HDFS (Hadoop Distributed File System), Hive, H-base, Pig, Chukwa, Avro, ZooKeeper etc. (2). We concentrate on developing a system that increases processing speed and capability to process huge amount of data. This contribution has an excellent demand in today's academic, medical, scientific and regular business. Document categorization is needed in the above fields; therefore it has a generalized application for commercial approach. We try to parallelize GA to improve the processing speed. We use Hadoops MapReduce and HDFS (Hadoop Distributed File System) framework approach. The rest of the paper is organized as follows. Basic System Architecture of the developed system is explained in section II. Section III explains Algorithm of system. Methodology used in our project is given in section IV. OlexGA is explained in section V. Section VI shows Results of our implementation. Concluding remarks and Future work are given in section VII.

Read the paper · More papers on PaperTik