Leveraging Topic Modeling and Extractive Summarization for Unlocking Insights from NeurIPS Papers
Sheikh Sharfuddin Mim, Doina Logofătu, Gabriel Guerrero-Contreras, Inmaculada Medina‐Bulo · 2024
Efficiently comprehending vast volumes of text, particularly in scientific literature, is a formidable challenge for researchers. To alleviate this burden, this paper investigates the application of topic modeling and extractive summarization techniques to analyze and distill valuable insights from scientific documents. Focusing on the NeurIPS dataset, which encompasses a diverse array of scientific papers, we employ techniques such as Latent Dirichlet Allocation (LDA) and probabilistic Latent Semantic Analysis (pLSA) to categorize documents based on their content. Subsequently, we implement a domain-specific extractive summarizer that leverages Convolutional Neural Networks (CNNs) for feature extraction and semantic segmentation. The summarizer is trained to identify and rank sentences based on relevance, using binarized ROUGE-2 scores as a metric. Our results demonstrate the effectiveness of these methods in condensing scientific documents, facilitating quicker information retrieval and understanding. The paper also compares the performance of LDA and pLSA, highlighting LDA’s superiority in topic coherence and generalizability. By leveraging these methods, we aim to demonstrate their efficacy in processing large volumes of text, facilitating enhanced information retrieval and understanding in the realm of scientific literature.