Text Vectorization Techniques for Trending Topic Clustering on Twitter: A Comparative Evaluation of TF-IDF, Doc2Vec, and Sentence-BERT

Alvian Daniel Susanto, Steven Andrian Pradita, Caroline Stryadhi, Karli Eka Setiawan, Muhammad Fikri Hasani · 2023

In this digital era, where technology is rapidly advancing, social media has become a primary platform for obtaining and disseminating information. Knowing what is being widely discussed and trending on social media is crucial for important aspects such as politics, economic, social and cultural issues. The objective of this research is to perform clustering on texts or sentences, and within each cluster, identify the most influential keywords that can serve as parameters to determine the topics being discussed in each cluster. Twitter was chosen as the social media platform to be analyzed in this research due to its text-based nature. The clustering method used is DBSCAN, considering that the number of clusters is unknown, and three text embedding techniques to be compared, namely TF-IDF, Doc2Vec, and Sentence-BERT. The performance of clustering with different text embedding techniques were evaluated using the silhouette coefficient. Hyperparameter tuning has been done to find the best-performing hyperparameters. From the best-performing technique, topic finding within the resulting clusters was conducted using Latent Dirichlet allocation (LDA). The results of this research indicated that clustering with DBSCAN and TF-IDF, with the highest silhouette coefficient, namely -0.00001, produced one cluster and 3342 outliers. DBSCAN and Doc2Vec, with the highest silhouette coefficient, namely 0.71590, produced one cluster and one outlier. DBSCAN and Sentence-BERT, with the highest silhouette coefficient, namely -0.02425, produced two clusters and two outliers. Based on the research findings, smaller silhouette scores tend to have a more varied number of clusters. DBSCAN with each tested text embeddings showed that the topic for every cluster, except for the first cluster of DBSCAN that use Sentence-BERT, were COVID-19 related topic. The DBSCAN and Sentence-BERT model, despite having a lower silhouette score, successfully identifies two separate clusters with distinct topics, whereas the other models only identify a single cluster.

Read the paper · More papers on PaperTik