METHODS & DESIGNS Generation of speech continua through monaural fusion
Dale C. Stevenson, John T. Hogan, Anton J. Rozsypal · 1985
A procedure based on monaural fusion has been developed to construct acoustic continua be tween natural speech sounds, to be used in studies of speech perception. Two speech stimuli of similar temporal structure and different spectral composition are precisely aligned in time and presented simultaneously to the listener. By mixing both stimulus components in varying inten sity ratios, a transition from one component to the other can be achieved. Such stimulus con tinua have several advantages over the synthetic continua commonly used in studies of categori cal perception and related phenomena: They are based on real speech stimuli; the endpoint stimuli are unambiguous; and the stimuli are characterized by a well-defined physical variable, the relative intensity of the two components. Traditionally, stimuli for recognition and discrimina tion experiments in speech research have been generated synthetically. Two endpoint stimuli representing the pro totypes for the response categories are synthesized, along with a series of intermediate stimuli whose parameters are interpolated between those of the two prototypes, thus creating a continuum. This article describes an alternative technique, monaural fusion, for constructing speech continua (Stevenson, 1979). The monaural fusion technique creates the stimulus continuum by aligning two speech stimuli of similar tem poral structure but different spectral composition and presenting them simultaneously to the subject. The two stimulus components ordinarily fuse to produce a single percept. When the relative intensities of the two com ponents are varied, a physical continuum that is categorically perceived is created. For instance, for syllable-initial stop consonants Ibl and Id/, the relative intensity continuum yields identification and discrimina tion results that are essentially equivalent to those obtained with traditional synthetic continua, obtained by inter polating formant frequencies (Repp, 1981; Stevenson, 1979).