Don’t make me guess : automatically detecting and naming topics in large collections of text

Andriy Kosar · 2025

With the rapidly increasing volume of information, it has become essential to categorize data for analysis and management. One technique that enables this is the use of “topics”: semantic shortcuts that let us label texts, group them by label, find relevant texts, or simply convey information about their content. Existing topic-detection methods, however, often fail to provide both reliable content categorization and human-friendly labels. Recent progress in topic detection, including improvements to methods based on Latent Dirichlet Allocation (LDA) with neural embeddings, semantic clustering with assigned top keywords, and transformer-based zero-shot classifiers shows promise but also has significant limitations. A key limitation is their limited world knowledge, which includes not only an understanding of real-world entities and facts, but also a grasp of the categories people commonly use to categorize information. The topic labels these methods produce are not always interpretable. In addition, they are often hard to adopt in dynamic environments because of their heavy computation and rigid frameworks. To overcome these challenges, this research introduces a novel dual approach that integrates predefined categorization with dynamic topic generation. The first approach leverages established taxonomies or internal taxonomies, utilizing a shared embedding space to transform both texts and known topics into a single semantic space. This enables unsupervised topical classification based on semantic similarity, eliminating the need for continuous model retraining as taxonomies evolve. It also allows for multi-label assignments through label-specific thresholding, ensuring that texts can belong to multiple relevant categories seamlessly. In addition to predefined categorization, the second approach employs Large Language Models (LLMs) to generate context-specific topic labels for individual texts. This research demonstrates that LLMs can generate topic labels comparable to those created by humans. This method easily identifies and names topics that may not fit into existing taxonomies. Based on the recognition of the subjective and contextual nature of how humans perceive and name topics, this research introduces a novel topic evaluation method based on information reconstruction, in which we assess the degree to which the assigned labels allow accurate reconstruction of the original text.

Read the paper · More papers on PaperTik