Developing AI that Pays Attention to Who You Want to Listen to: Deep-learning-based Selective Hearing with SpeakerBeam

Marc Delcroix, Tsubasa Ochiai, Hiroshi Satō, Yasunori Ohishi, Keisuke Kinoshita, Tomohiro Nakatani, Shoko Araki · NTT technical review · 2021

In a noisy environment such as a cocktail party, humans can focus on listening to a desired speaker, an ability known as selective hearing.In this article, we discuss approaches to achieve computational selective hearing.We first introduce SpeakerBeam, which is a neural-network-based method for extracting speech of a desired target speaker in a mixture of speakers, by exploiting a few seconds of pre-recorded audio data of the target speaker.We then present our recent research, which includes (1) the extension to multi-modal processing, in which we exploit video of the lip movements of the target speaker in addition to the audio pre-recording, (2) integration with automatic speech recognition, and (3) generalization to the extraction of arbitrary sounds.

Read the paper · More papers on PaperTik