Hybrid deep modelling with human knowledge in practical e-commerce search

Anxiang Zeng · 2023

In recent years, recommender systems have become increasingly important in the E-commerce marketplaces (e.g., Alibaba, Amazon).According to public information, recommender systems have contributed 30% gross merchandise value (GMV) for Amazon.Significant effort has been devoted into this research field by the industry and the academic community.Research works making key advances in recommendation systems can be divided into matrix factorization (MF), collaborative filtering (CF) and click-through rate (CTR) prediction.Currently, in industrial e-commerce recommendation systems, the main adopted approach is to use CF methods for performing recall tasks and deep learning-based CTR prediction methods for performing ranking tasks.However, significant challenges remain when these algorithms are to be deployed into practical e-commerce environments.In the system of search, most of the ctr modeling use big data, with a large number of user behavior data.The amount of data is a large number, which is very helpful to build the model.At the same time, these data also contains a lot of noise.This can lead to inaccurate modeling.In many cases, these behaviors are not consistent with business needs, we need to make manual guidance and intervention to the model.• Research Problem 1 -Classification of Long-Tail Queries: The first step of a typical e-commerce recommender system is query classification.Long-Tail Query in an e-commerce system typically refers to search queries that are not very common, which lacks of user behaviors(We define queries with a UV (Unique Visitors) of 100 or less in our system as long-tail queries).Existing works cannot account for the long-tail distribution of search queries, a key challenge that makes query classification more difficult than normal text classification tasks.Although there can be a few frequent queries that are xi xii searched thousands of times a day, most are infrequent.These infrequent queries make up a significant portion of the total query volume.In addition, many of them are short and ambiguous.Furthermore, the low frequency of long-tail queries makes them difficult to label.• Research Problem 2 -Evolving Multi-Objective Recommendation:There are often more than one objective an industrial recommender system needs to fulfill.The focus of research needs to shift from single objective optimization to multi-objective optimization that evolves over time in the face of changes in business requirements, especially in e-commerce promotion campaigns (e.g., Taobao's 11.11 Single's Day Shopping Festival).Dynamically adapting among the short-term and rapidly changing objectives is an important challenge as they can potentially conflict with each other.• Research Problem 3 -Cross-Domain Recommendation based on Instance Relationships: Item representation learning is crucial for search and recommendation tasks in e-commerce.In e-commerce, the instances (e.g., items, users) in different domains are always related.Such instance relationships across domains contain useful local information for transfer learning.However, existing transfer learning based approaches cannot leverage this knowledge.This thesis sets out to address these challenges to enable academic research in recommender systems to more readily benefit practical applications in the context of e-commerce: • For Research Problem 1: The category a query belongs to is closely related to different entities and their semantic roles in the query.Inspired by this observation, we proposed a novel model, E-BERT, which utilizes entity information to enhance the representations of long-tail queries and improve deep model performance on long-tail query classification.For each query, it recognizes the entities within it, and connects the query and entity with corresponding relations.In this way, a knowledge graph that contains different queries and entities can be constructed.During the optimization of query classification, an auxiliary knowledge graph embedding (KGE) module to learn entity knowledge from this graph.Additional entity knowledge is incorporated into the model by aligning the entity embeddings with the KGE xiii model pre-trained on Wikipedia data.The auxiliary module can be removed after training and, thus, incurs no extra computational cost.• For Research Problem 2: To allow dynamic switching among multiple business objectives to impact recommendation results, we proposed the online Deep Controllable Learning-To-Rank (DC-LTR) approach.It enhances the feedback controller in LTR with multi-objective optimization so as to maximize different objectives under constraints.It's ability to dynamically adapt to changing business objectives has resulted in significant business advantages.DC-LTR has become a core service enabling adaptive online training and real-time deployment ranking models for changing business objectives in AliExpress and Lazada.Under both everyday use scenarios and peak load scenarios during large promotional campaigns, DC-LTR has achieved significant improvements in adaptively satisfying real-world business objectives.• For Research Problem 3: To leverage instance relationships across recommendation domains, we proposed the Prior-Guided Transfer Learning (PGTL) approach.It utilizes such relationships to extract prior knowledge for the target domain and leverages it to guide the fine-grained transfer learning for e-commerce item representation learning tasks.Rather than directly transferring knowledge from the source domain to the target domain, the prior knowledge can serve as a bridge to link both domains and enhance knowledge transfer, especially when the domain distribution discrepancy is large.Since its deployment on the Taiwanese portal of Taobao in Aug 2020, PGTL has significantly improved the item exposure rate and item clickthrough rate compared to previous approaches.

Read the paper · More papers on PaperTik