On Clustering Algorithms: Applications in Word-Embedding Documents

Israel Mendonça · Journal of Computers · 2019

In this paper, we study the effectiveness of classical literature clustering algorithms applied to free text documents.We analyze the effects of varying the parameters on their performance and which aspects directly influence in the results.We apply a word-embedding-based technique to represent the document's bag-of-words and therefore be able to compare and study how these algorithms performs in the task of clustering these documents.We use two metrics that captures different aspects of the partitions and analyze those algorithms on the light of it.One of the main findings of this work is that some clustering algorithms are able to have a partition that's up to 91% of the real partition, whilst other performs really poor for the same dataset.We also find limitations on these techniques when trying to cluster hard datasets.

Read the paper · More papers on PaperTik