Jointly smoothing word embedding and text representation
Fatma Najar, Nizar Bouguila · 2021
Word embedding techniques have gained a lot of attention in recent years. Mapping words to distributions rather than points has tremendous advantages for language modeling and text documents classification. This paper presents a density word embedding based on categorical data distribution and Fisher information metric. Taking advantage of context information included in the embedded words, we derive a new latent topic model defined on a statistical manifold for the purpose of representing texts. We propose the mixture of smoothed Dirichlet distributions and we introduce the learning of mixture parameters using a likelihood-based approach. The proposed framework is validated through three gold-standard datasets: BBC News, Reuters 21578, and 20 Newsgroups. We compare the performance of the proposed approach with related works and the obtained results prove the robustness of our combined embedding approach with the smoothed latent topic model.