Deep neural network models of speech-in-noise perception for hearing technologies and research
Agudemu Borjigin, Kostas Kokkinakis, Hari Bharadwaj, Josh Stohl · The Journal of the Acoustical Society of America · 2022
Widespread adoption of artificial intelligence has yet to occur in the hearing field. Hearing technologies, such as cochlear implants (CIs), provide limited benefits of noise reduction, even with current state-of-the-art signal processing strategies. Recent developments in machine learning have produced deep neural network (DNN) models achieving remarkable performance in speech enhancement and source separation tasks. However, there are currently no commercially available CI audio processors that utilize DNN models for noise suppression. Furthermore, the current research community lacks a computational tool to match the complexity of natural auditory processing. To address these gaps, we implemented two DNN models: a recurrent neural network (RNN)—a lightweight template model for speech enhancement, and the SepFormer—the current top-performing speech-separation model in the literature. The DNN models resulted in significant improvements in terms of objective evaluation metrics, as well as intelligibility scores obtained with CI users at different signal-to-noise ratios. Given their flexibility and good performance on complex tasks, these models can also be used to generate hypotheses about speech-in-noise perception and serve as richer substitutes for models commonly used in research. This work serves as a proof-of-concept and a guide for the next steps towards integrating DNN technology into hearing technologies and research.