Embedding-Driven Clustering for Unerring Content Categorization in Low-Resource Hindi Language
Pushkar Baranwal, Amit Pundir, Sanjeev Singh, Geetika Jain Saxena · ACM Transactions on Asian and Low-Resource Language Information Processing · 2025
Content generation has been happening at a very high rate in many languages including the low resource local languages. It is crucial that this content is classified accurately to reach the right target audience within the shortest time possible. ML and NLP methods are employed for classification, organization, and storage of content for low-resource languages like Hindi. Implementation of clustering technique for categorizing similar articles into different groups tends to minimize the textual difficulty of handling large amounts of textual data. This helps ensure that the right news reaches the right person on time. This article focuses on grouping articles in Hindi using the unsupervised clustering such as FastText and other embeddings. Xlm-RoBERTa, IndicBERT, and LDA+IndicBERT. The Latent Dirichlet Allocation (LDA) coupled with IndicBERT has been found to be most efficient based on cluster performance, tested using the Silhouette and Davies Bouldin score. The overall certainty of LDA+IndicBERT is higher than other embedding techniques in terms of Silhouette score and the Davies Bouldin score and the cluster in word clouds and t-SNE plots. To enhance the interpretability of results, mainly for this cluster, we projected the embeddings in reduced-dimensions spaces using PCA and t-SNE. These visualizations provide additional insights into the structure of data and the dispersion of clusters. Themes of each cluster have been determined using word clouds in Hindi. Result shows LDA with IndicBERT embeddings for seven clusters yielded better Silhouette Score of 0.434 and Davies Bouldin score of 0.841, highlighting most distinct and well-defined clusters. Methodology validation has been done using other datasets which significantly shows the efficacy of clustering. This research paves the way for future research for clustering Hindi news articles, improving the organization and accessibility of news content for Hindi-speaking audiences.