Phaseper: A Complex-Valued Transformer for Automatic Speech Recognition

Delany Ramirez-Delrio, José A. Jaramillo-Villegas · IEEE Access · 2025

Recent advances in automatic speech recognition (ASR) have largely focused on real-valued neural network models (e.g., Whisper) that use only the magnitude of the speech signal’s spectrogram, discarding phase information. In this paper, we propose Phaseper, a novel complex-valued Transformer architecture for ASR that explicitly incorporates both magnitude and phase information of the speech signal. By processing complex-valued spectrogram inputs, Phaseper can leverage phase information that is usually ignored in conventional pipelines. We evaluate Phaseper on several speech recognition benchmarks, including LibriSpeech, Common Voice, TED-LIUM 3, and CHiME-6. Our experimental results demonstrate that Phaseper achieves competitive or improved word error rates (WERs) using a smaller training dataset compared to a real-valued Transformer baseline, with notable gains on datasets with noisy or reverberant audio (e.g., CHiME-6). Specifically, on the LibriSpeech dataset, Phaseper achieved a Word Error Rate (WER) of 0.54%, compared to real-valued neural network model’s WER of 2.7%. These findings highlight the importance of complex-valued representations in speech recognition.

Read the paper · More papers on PaperTik