Cross-Scale Attention for Long-Term Time Series Forecasting
Liangjian Wen, Quan Hu, Cong Guo, Ao Hu, Mingyi Zhang · IEEE Signal Processing Letters · 2024
Transformer-based models, especially PatchTST, have demonstrated remarkable success in time series forecasting tasks. However, the unique nature of time series data, which often contains jitter, and noise, and has inherently lower information density compared to images and texts, poses significant challenges. Specifically, the ViT-inspired patching design is suboptimal for time series data due to the sparse semantic relationships in such data. Moreover, modelling these sparse semantic relationships requires more resources and longer processing times. To overcome these limitations, we introduce a novel approach that leverages cross-scale attention interaction via a multi-scale patching technique. Our method initially treats the entire sequence as a single patch, then progressively divides it into increasing patches, doubling each time. This strategy improves the information density within each patch and reduces the total number of patches needed to model the time series to typically just seven effectively. We design a single layer of attention to model the cross-scale relationships among patches. These qualities significantly enhance computational efficiency. Extensive experiments have shown that our method surpasses or closely approaches existing methods in time-series forecasting benchmarks. Additionally, it achieves a speed that is 12x times faster than PatchTST on the large dataset.