Joint Iterative Multi-Speaker Identification and Source Separation using Expectation Propagation
John MacLaren Walsh, Youngmoo E. Kim, Travis M. Doll · 2007
The identification of individuals by the sound of their voices, a fairly straightforward task for humans, has proven to be quite difficult to achieve in a robust way computationally. The majority of past work in speaker (talker) identification has focused on the single speaker case, but these systems are easily confounded by most real-world settings where multiple talkers may be overlapping or speaking simultaneously. To address this situation, we propose a system that jointly identifies and separates the acoustic features of multiple talkers that fall within a library of known individuals. This system uses the probabilistic framework of expectation propagation (EP) to iteratively determine model-based statistics of both speaker identity and feature separation. This research has applications in audio surveillance as well as the forensic analysis of real-world sound recordings that contain multiple simultaneous talkers. Robust speaker identification could also lead to improved interfaces for human-computer interaction.