Enhancing Transformers without Self-supervised Learning: A Loss Landscape Perspective in Sequential Recommendation
Vivian W.-M. Lai, Huiyuan Chen, Chin‐Chia Michael Yeh, Minghua Xu, Yiwei Cai, Hao Yang · 2023
Transformer and its variants are a powerful class of architectures for sequential recommendation, owing to their ability of capturing a user’s dynamic interests from their past interactions. Despite their success, Transformer-based models often require the optimization of a large number of parameters, making them difficult to train from sparse data in sequential recommendation. To address the problem of data sparsity, previous studies have utilized self-supervised learning to enhance Transformers, such as pre-training embeddings from item attributes or contrastive data augmentations. However, these approaches encounter several training issues, including initialization sensitivity, manual data augmentations, and large batch-size memory bottlenecks.