TxtPrePro: Text Data Preprocessing Using Streamlit Technique for Text Analytics Process
Emad Qais, M. N. Veena · 2023
In natural language processing, cleaning up a lot of scraped text data is an important step that involves getting rid of irrelevant and noisy data from the text corpus. Text data obtained from web scraping or other sources may contain various unwanted elements such as HTML tags, non-textual characters, and punctuation marks, which can negatively impact the accuracy and efficiency of NLP algorithms. The study simplifies obtaining text data from multiple websites by employing TxtPrePro, a simple pipeline for scraping and text preprocessing that can be used for topic modeling, text summarization, sentiment analysis and other purposes. The proposed method used web scraping to collect a large amount of text data then performed proper text preprocessing, obtained some information such as tables visualizations and obtained a cleaned corpus.