Text Processing and Analysis Pipeline for Scientific Literature
Inderjeet Singh, Satyam, Ashutosh Semwal, Shivansh Singh, Gouri Gupta · 2024
This paper presents an exploratory data analysis-driven text processing pipeline for scientific publications based on the Scopus dataset. Sorting texts, spotting trends, and making recommendations are the goals. Pre-processing, cluster, and classification analyses are all incorporated into the suggested methodology, which produces a user recommendation software system. A parametrical method that takes user preferences into account extracts semantic information to overcome the obstacles associated with data preparation. An ensemble approach is used for user profile clustering, and an entropy-based ensemble technique is used for classification. The NLoN software performs exceptionally well when it comes to textual data classification. The analysis of the Scopus dataset highlights the importance of data processing pipelines by assisting with entity extraction and sentence classification. By combining information retrieval, data mining, and natural language processing, the work improves text processing for scientific publications.