Integrating TF‐IDF Features to Divide Amazon Product Reviews into Positive and Negative Groups
Ankit More, Abhishek Mishra, Prakash Maravi, Prathamesh Muzumdar, Abhishek Sharma · 2024
Many real-world applications benefit from text mining and similar techniques, such as customer management in business intelligence systems, retrieving medical data and using language analysis to identify fraud. There is usage of natural language processing and data mining. In these applications to help process the data and extract the necessary patterns. Text mining algorithms are employed in this work to classify product reviews on Amazon. This analysis can be useful for consumers in helping them decide what to buy, and by providing feedback on the product, it might also help the person who produces it. The dataset of Amazon product reviews originating from Kaggle is utilized in this instance. In order to remove stop words and special characters from the text data, the data set is first preprocessed. The second step after that several feature selection methods have been used. Possible keywords are selected from the reviews using the Part of Speech Tagging (POS) method based on Natural Language Processing (NLP) parser, and features are extracted using the Term Frequency-Inverted Document Frequency (TF-IDF) based feature selection methodology. Following the combination of the two features, training and testing datasets are created using the combined data features. The Support Vector Machine (SVM) is then trained on the training dataset, and the test dataset that is produced is used for validation. The results of the training process. As sample sizes increase, investigations are conducted using a range of sample sizes. Furthermore, measurements have been made of the accuracy, memory performance, time, recall, and F1-score. The model's performance demonstrates improved accuracy and less resource use. Finally, a few more developments of the work are also recommended in light of the experimental research.