The separation of speech from interfering sounds: an oscillatory correlation approach
Guy J. Brown, DeLiang Wang · 2003
A neural model is described which uses oscillatory correlation to segregate speech from interfering sound sources. The core of the model is a two-layer neural oscillator network. The first layer of the network identifies the connected regions of energy in the time-frequency plane (segments). In the second layer, segments that have a common fundamental frequency are grouped into streams. A stream is represented by a synchronized population of relaxation oscillators, and different streams are represented by desynchronized oscillator populations. The model has been evaluated using a corpus of voiced speech mixed with interfering sounds, and produces an improvement in signal-to-noise ratio for every mixture.