Multilingual Text Clustering and Labeling Using OpenAI Embedding

Yoshihiro Adachi, Minoru Uehara · 2025

The OpenAI embedding model is highly multilingual. Embeddings of words and sentences with similar meanings written in different languages generated by this model show high cosine similarity. We applied this finding to implement a multilingual topic analysis system using OpenAI embeddings. The system clusters a collection of English, Chinese, and Japanese sentences into groups based on topic similarity and labels each cluster with a list of words that clearly describe the topic characteristics of the cluster in each language. We developed a method to evaluate the appropriateness of cluster labels based on the reproducibility of the clusters. Evaluation using this method showed that it is appropriate to select the words that make up each cluster label based on their proximity to the cluster centroid.

Read the paper · More papers on PaperTik