Leveraging NLP and Deep Learning for Automating Product Categorization

Kalyana Kumar P, Mehta Anamika, S Snehaa · 2024

Classification of products in a specified category is key to effective stockkeeping, better systems that suggest products to buyers, and an overall more pleasant shopping experience when one shops online. Attempts to address the research problem in practice with the use traditional methods such as rule based systems, manual tagging or even simpler machine learning models like Support Vector Machines (SVM) are usually proven to be ineffective, inefficient and non-scalable in nature. The paper proposes a system that employs machine learning for product category classification. Standard Natural Language Processing techniques such as tokenization, stemming, stop word removal as well as vectorization techniques such as Count Vectorizer are applied on the product descriptions. For classification purposes, BERT based Contextual Embeddings and a mutually inclusive Random Forest (RF) and Logistic Regression (LR) based model have been employed in the project to ensure correct category creation with high accuracy. As a data set includes a variety of product descriptions collected from different e-commerce sites which indicates the real-world complexity and variety of data. Key performance indicators like accuracy, precision, recall and F1-score are useful in training the model. Results indicates that performance of BERT-based deep learning network architecture with hybrid RF + LR outshines all the existing methodologies including SVM. Also the Project assesses how data augmentation, transfer learning and model tuning affect classification accuracy. The results also emphasize how NLP and deep learning can scale up and be effective in automating product categorizations. The future improvement will be integrating more sophisticated techniques in NLP and multi-modal data and also user feedback in order to improve the performance of the system even further.

Read the paper · More papers on PaperTik