Improving Arabic document clustering using K-means algorithm and Particle Swarm Optimization
Abdullah S. Daoud, Ahmed Sallam, Mohamed E. Wheed · 2017 Intelligent Systems Conference (IntelliSys) · 2017
Document clustering plays a vital role in text mining fields such as information retrieval, sentiment analysis, and text organizing. Document clustering aims to automatically divide a collection of documents based on some aspects of similarity into groups that are meaningful, useful or both. This paper aims to improve the clustering task for the Arabic documents. Recent studies show that partitioning clustering algorithms are more suitable for clustering process. However, k-means is the most common algorithm that is being used for clustering process because of its simplicity and speed. It can only generate an arbitrary solution because the results depend on the initial centers for the desired clusters “the seeds”. In this paper, a new modified k-means algorithm called PSO K-means, supported by Particle Swarm Optimization (PSO) is applied to enhance the Arabic document clustering process. Then, an intensive comparative study between the proposed model and the standard k-means algorithm is applied. Also, the stemming algorithms those are being used in Arabic language processing were assessed. Through the experiments, an evaluation for the new algorithm is done with three different Arabic data sets. The results demonstrate that the proposed model can produce more accurate results compared to the standard k-means algorithm for Arabic language documents. On the other hand, Arabic light stemmer is more suitable for the stemming step.