SPATIOTEMPORAL TRANSFORMER FRAMEWORK FOR NEXT-GENERATION WIRELESS CHANNEL ESTIMATION
L. Balaji, A. Dhanalakshmi, B Chempavathy, S. Aswini, S Zoya Parvez, D Sridhar · Telecommunications and Radio Engineering · 2026
The work proposes an architectural approach called dual-stream transformer for predicting channel state information (CSI) in highly mobile millimeter wave (mmWave) massive MIMO (multiple in-put/multiple output) systems. This method includes two streams: a spatiotemporal transformer (ST-transformer) for modeling temporal trends and spatial relationships between CSI matrices; and a complex-valued transformer (CV-transformer) that models complex-valued representations to capture amplitude-phase relationships. These two streams are combined with a cross-attention fusion mechanism to generate accurate CSI reconstructions. Evaluation of this architecture was completed via the use of synthetic data sets that were created by employing the 3GPP TR 38.901 channel model for CSI generation under a variety of conditions, including 30-100 km/h mobility speeds and 0-20 dB signal-to-noise ratio (SNR) levels. Results indicated that this dual-stream transformer-based architecture provided better performance than the conventional minimum mean square error (MMSE), long shortterm memory (LSTM), and single-stream transformer architectures. In terms of CSI normalized mean square error (NMSE), the dual-stream transformer architecture resulted in NMSE values of -15.2, -13.7, and -11.4 dB at 30, 60, and 100 km/h, respectively. Spectral efficiency values were also calculated as 1.29, 3.11, and 5.43 bits per second per Hz (bps/Hz) at 0, 10, and 20 dB SNR, respectively. Finally, despite its dual-stream design, the authors indicate that the architecture has an inference time of approximately 4.88 ms/frame, which is sufficiently low to support real-time CSI estimation in 5G-Advanced and 6G wireless networks.