Common-Gain Autoencoder Network for Binaural Speech Enhancement

Stefan Thaleiser, Gerald Enzner, Rainer Martin, Aleksej Chinaev · IEEE Open Journal of Signal Processing · 2025

Binaural processing is becoming an important feature of high-end commercial headsets and hearing aids. Speech enhancement with binaural output requires adequate treatment of spatial cues in addition to desirable noise reduction and simultaneous speech preservation. Binaural speech enhancement was traditionally approached with model-based statistical signal processing, where the principle of common-gain filtering with identical treatment of left- and right-ear signals has been designed to achieve enhancement constrained by strict binaural cue preservation. However, model-based approaches may also be instructive for the design of modern deep learning architectures. In this article, the common-gain paradigm is therefore embedded into an artificial neural network approach. In order to maintain the desired common-gain property end-to-end, we derive the requirements for compressed feature formation and data normalization. Binaural experiments with moderate-sized artificial neural networks demonstrate the superiority of the proposed common-gain autoencoder network over model-based processing and related unconstrained network architectures for anechoic and reverberant noisy speech in terms of segmental SNR, binaural perception-based metrics MBSTOI, better-ear HASQI, and a listening experiment.

Read the paper · More papers on PaperTik