Generate Topics and Filter Novelties at Scale with Pre-Trained Language Models: A GPT-2 Method
B. Jeevashri, Uma Priyadarsini P. S · 2025
The exponential rise of digital material calls for the development of more sophisticated approaches to the formation of topics that guarantee semantic richness, diversity, and scalability within the content. This research introduces an innovative pipeline employing the pre-trained Generative Pre-trained Transformer 2 (GPT-2) model to tackle issues related to topic coherence, redundancy, and the effective processing of extensive textual datasets to integrates text segmentation, novelty filtration, and semantic coherence enhancement. Text chunking facilitates the processing of extensive datasets by partitioning information into manageable chunks, whereas novelty filtering utilizes Term Frequency-Inverse Document Frequency (TF-IDF) vectorization and cosine similarity measures to remove redundancy. Experiments on a news dataset illustrated the pipeline's efficacy, with filtered GPT-2 outputs attaining higher coherence scores (0.85) than raw GPT-2 (0.75) and conventional approaches such as Latent Dirichlet Allocation (LDA) (0.45). Chunking was essential for sustaining scalability, with an ideal size of 800 tokens achieving a compromise between computational efficiency and precision. The methodology substantially surpassed traditional methodologies in producing distinctive, contextually relevant themes, rendering it appropriate for applications including content analysis, document organization, and knowledge discovery. Future endeavors will concentrate on optimizing for domain-specific datasets, multilingual applications, and minimizing computational overhead to improve adaptability and efficiency.