Customized Transformer for Confident Browse Node Classification on Imbalanced Product Data

Kumar Selvakumaran, Aman Sami, Benazir Begum · 2023

Rapid digitization of the world has revitalized the potential of AI. As a result, this phenomenon awakened a huge demand for data-mining pipelines that can efficiently process massive volumes of data to generate valuable insights. For instance, the e-commerce industry is a hub for such high-capacity analytical pipelines as the big players in this industry deal with data pertaining to millions of products, transactions, and users. Consequently, the sheer amount of data used in such AI workflows gives rise to Scalability issues and resource-efficiency issues. As a result, there is a necessity for more computationally efficient approaches that optimize resource allocation and minimize infrastructure costs. Thus, this paper focuses on implementing an efficient product classification pipeline using a customized version of a state-of-the-art transformer-based model which is able to efficiently process data volumes that are comparable to the current voluminosity of industrial data. This customized workflow adapts a loss function that was originally developed for object detection and this adaptation has proved to help the model converge faster and combat class imbalance which is a serious issue in developing large-scale Natural Language Processing systems.

Read the paper · More papers on PaperTik