Inplace Frequency Filtering and Cepstral Speech Modeling in Binaural Speech Enhancement

Jinjiang Liu, Hao Li, Fei Chen, Zhiyong Wu, Xueliang Zhang · IEEE Transactions on Audio Speech and Language Processing · 2025

This paper presents an improved Inplace Cepstral Convolutional Recurrent Neural Network (ICCRN+) model and evaluates it on the binaural speech enhancement task. The proposed ICCRN+ offers the following advantages: 1) By removing spectral downsampling, inplace neural networks preserve spatial cues and reverberation details as much as possible for better spatial filtering, dereverberation, and binaural cue preservation; 2) The neural cepstral modeling technique directly transforms frequency-domain neural network features to a cepstral space, thereby achieving effective harmonic restoration. This transformation utilizes the fast Fourier transform (FFT) to achieve full-band feature perception with$O(n\log {n})$efficiency. 3) The consistent feature shape in inplace neural networks provides extra freedom to explore novel and efficient network topology designs in this paper. Comprehensive evaluations on dereverberation and denoising task demonstrate ICCRN+'s significant improvements in speech intelligibility and quality over baseline models, as well as the unique advantages of inplace neural networks in binaural cue preservation.

Read the paper · More papers on PaperTik