A simulation of vowel segregation based on across-channel glottal-pulse synchrony
Daniel P. W. Ellis · 1994
As part of the broader question of how it is that human listeners can be so successful at extracting a single voice of interest in the most adverse noise conditions (the ‘cocktail-party effect’), a great deal of attention has been focused on the problem of separating simultaneously presented vowels, primarily by exploiting assumed differences in fundamental frequency (f0) (see (de Cheveigné, 1993) for a review). While acknowledging the very good agreement with experimental data achieved by some of this models (e.g. Meddis & Hewitt, 1992), we propose a different mechanism that does not rely on the different period of the two voices, but rather on the assumption that, in the majority of cases, their glottal pitch pulses will occur at distinct instants. A modification of the Meddis & Hewitt model is proposed that segregates the regions of spectral dominance of the different vowels by detecting their synchronization to a common underlying glottal pulse train, as will be the case for each distinct human voice. Although phase dispersion from numerous sources complicates this approach, our results show that with suitable integration across time, it is possible to separate vowels on this basis alone. The possible advantages of such a mechanism include its ability to exploit the period fluctuations due to frequency modulation and jitter in order to separate voices whose f0s may otherwise be close and difficult to distinguish. Since small amounts of modulation do indeed improve the prominence of voices (McAdams, 1989), we suggest that human listeners may be employing something akin to this strategy when pitch-based cues are absent or ambiguous. 1.