Automated two speaker separation system
Kwang-Shik Min, D. Chien, Shiyou Li, C. Jones · 2003
An algorithm designed to improve results for separating two voices simultaneously recorded on a single channel is presented. A variable frame size orthogonal transform and a spectral matching technique are used. A multistep pitch detection scheme is proposed which includes a traditional autocorrelation function, a modified autocorrelation, the average magnitude difference function, and a look-forward and look-backward double checking scheme. The orthogonal transforms utilized include the fast Fourier transform and the fast triangular transform. For a variable frame size transform, a prime factor fast Fourier transform has been developed. The execution of the process is automated and implemented on the IBM-PC, VAX 8650, and HP 9000. Intelligibility tests using simple quantitative measures have been performed on the separated signals. An extension of the problem to the three-speaker case is reported.>