VISUALIZING DOCUMENT CLUSTERS WITH BERT, K-MEANS, AND DBSCAN

D.P.V.Phani Rajakumar, Dama Manish Kumar, Chiruvolulanka Mohith Nancharaiah, Dusanapudi Lakshmi Naga Venkata Likitha, Daravath Veera Prasad Nayak, Boppana Rishitha · Journal of Nonlinear Analysis and Optimization Theory & Applications (JNAO) · 2025

The rapid growth of digital repositories containing textual data, such as research articles, news stories, and reviews, necessitates effective clustering techniques for categorization and information retrieval. However, traditional clustering methods like KMeans and DBSCAN often struggle with high-dimensional, sparse text data. This paper proposes an approach that leverages BERT embeddings to enhance clustering performance. By integrating BERT embeddings with K-Means and DBSCAN, we improve clustering accuracy by capturing the semantic richness of textual data. Experimental results on the BBC Full Text dataset demonstrate superior clustering performance, achieving a 91% accuracy based on Purity and Adjusted Rand Index (ARI). Furthermore, an interactive visualization component is introduced to aid in the interpretation of clustered data.

Read the paper · More papers on PaperTik