A Hybrid Algorithm for Data Clustering Using Honey Bee Algorithm, Genetic Algorithm and K-Means Method

Mohammad Ali Shafia, Mohammad Rahimi Moghaddam, Rozita Tavakolian · 2011

With the rapid increasing of information on the web, clustering related data and documents to achieve useful information will be more important for information retrieval systems. There exist a lot of methods to tackle the issue of clustering. The most important of which is K-Means with classifying the data set into a number of homogenized groups based on their similarities. The main problem with K-Means is the tendency to convergence in Local Optimum. Metaheuristic algorithms are widely used to optimize the result of K-Means. But, resulting clustering quality and the optimal solution is still a challenging task. Honey Bee Algorithm, as a Metaheuristic Algorithms, introduces a novel approach to search in solution space via observing honey bee behavior regarding foraging that leads to improving solution quality. In this article a novel population based hybrid algorithm called GBTKC is developed based on basic Honey Bee Algorithm in which the benefits of K-Means Method is used in order to improve its efficiency. Also, the simplicity of K-Means, the diversity of Genetic Algorithm to find the global optimum and advantages of Tabu Search has been combined in GBTKC. Using Honey Bee, our hybrid algorithm has more ability to search for the global optimum solutions and more ability for passing local optimum as well as generating efficient near optimal solutions. Moreover, GBTKC is run on three known data sets from the UCI Machine Learning Repository and the results of clustering using this algorithm compared with other studied algorithms will be stated in the literature review. So, this is revealed that GBTKC is definitely a convergent optimal solution, and the quality of answers provided by the algorithm is more reliable than those of other algorithms.

Read the paper · More papers on PaperTik