Bengali Document Clustering Using Word Movers Distance
Adnan Ahmad, Md. Ruhul Amin, Farida Chowdhury · 2018
In this paper, we propose a pipeline architecture for Bengali Document clustering and apply it for clustering Bengali news documents using different clustering algorithms. Our goal is to cluster news from different online newspapers according to the topic describing the identical stories. We used Word Movers Distance (WMD), a relatively new algorithm to measure document distances, which is based on vector representation of words. Later, we conduct Bengali document clustering using several other algorithms, namely K-means, Hierarchical Clustering Algorithm (HCA) and Hierarchical Density-Based Spatial Clustering of Applications with Noise (HDBSCAN). We also evaluate the clusters against a manually prepared set of clusters, which we consider as the ground truth. Our experiment shows that, HCA performs best with a Fl-score of 92%, which is the most similar to the number of clusters and cluster members compared to the ground truth. We also released a live working demo where the program collects the recent news from popular online Bengali newspapers and create clusters according to the news story.