Deep Learning Based Distant-talking Speech Processing in Real-world Sound Environments
Shoko Araki, Masakiyo Fujimoto, Takuya Yoshioka, Marc Delcroix, Miquel Espi, Tomohiro Nakatani · NTT technical review · 2015
This article introduces advances in speech recognition and speech enhancement techniques with deep learning.Voice interfaces have recently become widespread.However, their performance degrades when they are used in real-world sound environments, for example, in noisy environments or when the speaker is some distance from the microphone.To achieve robust speech recognition in such situations, we must make progress in further developing various speech processing techniques.Deep learning based speech processing techniques are promising for expanding the usability of a voice interface in real and noisy daily environments.