Recurrent Neural Beamformer for Multichannel Speech Enhancement Under Adverse Noise Condition
Zhi-Wei Tan, Yuan Liu, Andy W. H. Khong, Anh H. T. Nguyen · IEEE Transactions on Audio Speech and Language Processing · 2025
Speech enhancement under low signal-to-noise ratio conditions, such as for UAV applications, is challenging. Existing deep learning with beamforming approaches requires prior direction information or a sufficiently large number of estimated speech frames. We propose a recurrent neural beamformer (R-NBF) to enhance multichannel speech signals. The proposed R-NBF architecture comprises a set of beamformers followed by a multichannel complex spectral mapping twin-dilated U-net (TDU-net) and a feedback connection. This feedback connection serves as a conduit that facilitates frame-based adaptation between the output of the TDU net and the set of beamformers. We present an analysis framework based on Taylor's first-order approximation and Wirtinger's calculus that illustrates how the iterative optimization process results in a gradient that reduces the spatial covariance matrix error. The proposed approach is evaluated via simulation and on data recorded from a hexacopter hovering above an open field.