An Intra- and Inter-Frame Sequence Model with Discrete Cosine Transform for Streaming Speech Enhancement

Yuewei Zhang, Huanbin Zou, Jie Zhu · 2024

Nowadays, in order to improve the speech enhancement performance, many methods attempt to reconstruct the target magnitude and phase spectrum simultaneously. They usually process the complex short-time Fourier transform (STFT) spectrum, leading to a huge model complexity. In this paper, we utilize the short-time discrete cosine transform (STDCT) rather than STFT. Since STDCT is a lossless real-valued transformation with implicit phase, our method achieves an excellent performance with lower complexity. Besides, we take convolutional recurrent network (CRN) as the network backbone, and design a dual sequence modeling block to capture the intra-frame correlation among different frequency bins and the inter-frame context along the time dimension simultane-ously, so we name our model IICRN. The experimental results indicate that IICRN achieves superior performance over previous advanced methods.

Read the paper · More papers on PaperTik