Arabic words clustering by using K-means algorithm

Dheyaa Shaheed Al-Azzawi, Faiez Musa Lahmood Al-Rufaye · 2017

These We mean by the clustering is a technique to divide the text into clusters of words, so that words in the same cluster are similar to each other. As humans we can have significant difficulty understanding the clustering of words to each other. To a machine, it represents a huge challenge. To address this problem, this paper describes a new words-clustering technique based on certain text characteristics; by building a system to cluster words in the text depending on characteristics such as morphological, syntactic and Semantic. The clustering is a method of Unsupervised Machine Learning methods, where it collects words with other have similar characteristics in the clusters based on Similarity Function to calculate the distance between those words. We depended on k-mean Clustering to calculate the distance between words, then generating clusters for all referred words in the text. Finally, we will evaluate our system results by using the common evaluation methods. There are three methods of evaluation methods: Precision, Recall, and F-Measure.

Read the paper · More papers on PaperTik