Exploring Pre-processing Strategies and Feature Extraction in practical aspect for Effective Spam Detection
Harjeet Singh, Shivani Sood, Heranmoy Maity, Yogesh Kumar · 2024
With the advent of Large Language Models (LLMs) based ChatGPT, a drastically change has been observed in the areas of Natural Language Processing. LLMs have shown remarkable performance in various natural language processing tasks and applications. This article first presents all the stages (roadmap) towards implementing LLMs, wherein pre-requisite tools and methods for constructing a Large Language Model are discussed. In this connection, required text data pre-processing approaches have been disscused and implemented using the publicly, online available dataset (SMSSpamCollection). By using this dataset,two classes ("ham" and "spam") classification problem has been addressed. In this classification, we first elaborate the working of two text feature extraction techniques, namely, Bag-of-Words (BOW) and Term Frequency-Inverse Document Frequency (TF-IDF) and then train the classifier by using both the features. Testing experiments produced an appreciable accuracy rate 97.22% and 98.47% for BOW and TF-IDF, respectively. Although, classification accuracy is good enough but these extracted features have certain limitations, which are also discussed in this article. In the future, we will explore this work with various neural network architectures, commonly used in LLMs in order to overcome these issues.