LLM-Based Topic Modeling for Dark Web Q&A Forums: A Comparative Analysis With Traditional Methods

Luis de‐Marcos, Adrián Domínguez‐Díaz · IEEE Access · 2025

Topic modeling is a critical tool for understanding the thematic structures of unstructured text data, particularly in specialized domains like the dark web. This study compares the effectiveness of large language models (LLMs) and traditional topic modeling techniques in analyzing dark web Q&A forums, which are characterized by short, informal, and context-specific posts. We evaluate two LLMs—GPT and Gemini—against traditional methods, TF-IDF (scikit-learn) and Latent Dirichlet Allocation (LDA, Gensim), in their ability to generate meaningful and coherent topics. Additionally, we explore the impact of topic granularity by comparing GPT single-word topics with two-word topics. Using semantic similarity and Levenshtein distance metrics, we quantify the alignment and divergence between the topic representations produced by these methods. Our findings demonstrate that LLMs consistently outperform traditional methods in capturing contextually relevant themes, such as "scam," "bitcoin," and "hacking," while traditional techniques often produce generic or non-topical terms like "questions" and "know." Semantic similarity and Levenshtein distance metrics further highlight the strong alignment between LLMs and the divergence between LLMs and traditional methods. The comparison between single-word and two-word topics reveals that while two-word labels offer additional nuance, their benefits are limited in the context of dark web forums, where concise posts often make single-word topics sufficient. These results underscore the importance of selecting appropriate topic modeling techniques based on the characteristics of the text and the requirements of the analysis. This study highlights their potential as a valuable tool for analyzing complex and unstructured text data, particularly in specialized domains like the dark web.

Read the paper · More papers on PaperTik