Multi-view Topic Modeling Using Multi-text Representations

S. Abilasha, Rafika Boutalbi, Stéphane Delliaux · 2025

Topic modeling has been widely used for decades across various applications, particularly in web data analysis. The Latent Dirichlet Allocation (LDA) is based on statistical modeling, where words are modeled as a distribution of words, and documents as a distribution of topics. Today, there are several text ways to represent texts, from the simplest one Bag-of-Words to the recent word embedding. Two main classes of word embeddings exist, static word embeddings (Word2Vec, Glove, etc) and contextual word embedding (BERT, RoBERTa, etc). And, some recent works, showed the interest and the effectiveness of integrating such text representations to improve topic modeling performance. However, in the unsupervised context of topic modeling, there is no prior knowledge that can ensure the effectiveness of a specific text representation. Also, they often face the challenge of topic collapsing, where identified topics become semantically redundant, resulting in overly similar topics, limited topic diversity, and reduced interpretability of the model. In this paper, we propose a new topic modeling approach that can consider multi-text representations simultaneously named Multi-view Topic Modeling (MTM). The MTM algorithm is able to discover topics based on several text representations and generate multi-view topic embeddings allowing us to interpret the results considering several views. We evaluate the MTM algorithm on five real-world datasets and show the effectiveness of the proposed algorithm in terms of topic coherences and clustering.

Read the paper · More papers on PaperTik