Inferring the Number and Order of Embedded Topics Across Documents
Asana Neishabouri, Michel C. Desmarais · Procedia Computer Science · 2021
Documents are often organized according to an embedded structure, where a set of documents covers a topic and gets extended to a more specialized topic. We refer to this structure as embedded topics and address the issue of inferring the number and order of topics in a given corpus. While this problem is akin to finding clusters of documents and has been addressed in numerous studies in areas such as topic modeling, information extraction and knowledge discovery, we show that existing approaches are not effective in the specific context of embedded topic structures, and propose a novel technique for that purpose. We also propose an approach to uncover the order of such embedded topics. To determine the number of topics, the proposed method relies on the analysis of eigenvalues of a conditional probability matrix derived from the document-term matrix. We use Kmeans to determine the actual topic clusters, and conditional probability computation to determine the order. We compare the performance of our method to alternative methods for determining clusters and dimensionality. Results show that the proposed approach can effectively derive the right number of topics and embedding structure order.