Topic-Based Clustering of Japanese Sentences Using Sentence-BERT

Kenshin Tsumuraya, Miki Amano, Minoru Uehara, Yoshihiro Adachi · 2022

In recent years, text analysis and generation techniques using machine learning have substantially developed and are used to analyze opinions on social networking services and reviews on the Web. However, to build a supervised learning model with BERT to classify sentences based on topics, it is necessary to select topic classes in the field of application and create teacher-labeled datasets corresponding to each class. In this study, we developed a method for clustering Japanese sentences based on topics using a distributed representation obtained using Japanese Sentence-BERT (JSBERT) fine-tuned by the Japanese translation of the Stanford Natural Language Inference corpus. In particular, the use of a distributed representation generated by JSBERT only from the nouns that make up a sentence is an effective way to cluster Japanese sentence datasets based on topics using cosine similarity. We also devised a method to assign an appropriate cluster label to each cluster to make it easier to understand the contents of the cluster. Furthermore, we devised a function to explain why each sentence was classified into the corresponding cluster.

Read the paper · More papers on PaperTik