SpeakerBeam: A New Deep Learning Technology for Extracting Speech of a Target Speaker Based on the Speaker's Voice Characteristics

Marc Delcroix, Kateřina Žmolíková, Keisuke Kinoshita, Shoko Araki, Atsunori Ogawa, Tomohiro Nakatani · NTT technical review · 2018

In a noisy environment such as a cocktail party, humans can focus on listening to a desired speaker, an ability known as selective hearing.Current approaches developed to realize computational selective hearing require knowing the position of the target speaker, which limits their practical usage.This article introduces SpeakerBeam, a deep learning based approach for computational selective hearing based on the characteristics of the target speaker's voice.SpeakerBeam requires only a small amount of speech data from the target speaker to compute his/her voice characteristics.It can then extract the speech of that speaker regardless of his/her position or the number of speakers talking in the background.

Read the paper · More papers on PaperTik