Pronunciation variation within and across speakers
Michael H. Cohen, Jared Bernstein, Hy Murveit · The Journal of the Acoustical Society of America · 1987
An understanding of the structure of pronunciation variation over a population of speakers and over time in the utterances of one speaker should be useful in designing speaker-independent speech recognizers. This paper reports a series of experiments designed to show different kinds of patterns observed in alternative forms of words in constant contexts (e.g., the presence or absence of frication in the “y” in “had your”). An analysis is presented of transcribed data from 630 speakers reading two sample sentences as well as data from four speakers reading the same two sentences 24 times each, separated by filler material, in three separate sessions. The analysis quantifies the relative usefulness of competing models of variation in information theoretic terms. The results indicate that (1) speakers can be clustered into low variation groups such that the variation within a group is significantly less than the population variation, and (2) individual speakers show greater consistency than comparable clustered subsets of the population. Finally, it is suggested how this structure may be used to guide rapid, automatic adaptation in speech recognition. [Work supported by NSE.]