Efficient Wavelet Attention with Trainable Frequency Filter for Multi-modal Sequential Recommendation
Yuang Jia, Ruiting Dai, Lisi Mo, Ke Qin · 2025
In recent years, sequential recommendation systems have focused on modeling users' historical interactions to capture their dynamic preferences, thereby achieving more accurate and personalized recommendations. Current approaches typically utilize user and item IDs along with textual features for sequence modeling, employing self-attention mechanisms or Fourier transforms to capture long-range dependencies and extract sequence features. Despite their achievements, these methods still suffer from overfitting and transferring knowledge to new datasets. The issue lies in the lack of effective inductive biases in self-attention methods and the limitation of Fourier transforms, which only capture frequency-domain features, restricting the comprehensive modeling of users' dynamic behavior across both time and frequency domains. To tackle these issues, we propose a novel Wavelet attention-based Time-frequency domain Multi-modal Sequential Recommendation model (WTMSRec). WTMSRec consists of three core components: efficient data extraction and augmentation, a learnable frequency filter, and an innovative wavelet attention mechanism. It operates autonomously without reliance on user or item IDs, proficiently capturing and generalizing dynamic variations in user interests across both time and frequency domains. Experimental findings on five distinct downstream datasets validate that WTMSRec surpasses the performance of current cutting-edge models, with notably enhanced training efficiency.