THE INFLUENCE OF SPEAKER NORMALISATION AND TRANSMISSION CHANNEL COMPENSATION ON VOWEL IDENTIFICATION IN NATURAL SPEECH

AJ WATKINS, SK BOEGLI · 2024

lnfonnation beyond the syllable appears not to be essential for vowel identi cation as listeners are good at identifying isolated vowels [1].However other experiments suggest that this 'syllable extrinsic ' [2] information does make some contribution.The influence of a carrier phrase on a test sound has been widely investigated.The idea is that with different talkers different reference frames are utilised in vowel identi cation [3].When shifts in the first formant of a carrier phrase signal a different speaker, the perceptual identity of a subsequent vowel is changed [4,5].However, equating the long term average spectra of formant-shifted carriers eliminates these perceptual effects, thus a compensation for averaged spectral characteristics is sufficient to account for these types of 'speaker normalisation' result [5]» Long term spectrum compensation is normally associated with adjustments for the spectral envelope distortion when a sound is transmitted from a source to a listener [6,7].Effects of 'syllable extrinsic' information on vowel quality.have been attributed to the use of synthetic speech.However, some experiments have used natural voices, nding that the identity of a male speaker s vowel can be changed by embedding it in a child s carrier [8,9].Here we use natural voices and attempt to change the identity of a male and a female speaker's vowels by embedding them in each others sentences (experiment 1).Experiment 2investigates whether equalising the long term average spectra of these carrier sentences eliminates perceptual changes to the vowels.this is an attempt to replicate the synthetic speech finding with natural speech.The third experiment uses reversed carriers to ask whether phonetic or auditory mechanisms are at work: a reversed speech carrier does not conform to the constraints of natural utterances and might reduce any effects of a phonetic mechanism. GENERAL METHODA male and female were recorded rapidly saying various lth/ test words in the sentence frame "Please say for me .The male imitated the female in speaking rate and stress pattent.Two types of recordings were made {or the male's carrier sentence.one spoken in his normal pitch and one in which he mimicked the female's pitch.Mixed speaker sentences were created by embedding the male test words in the female speaker's carrier sentence, and by embedding female tcst words in both the original male and the male imitating female carrier sentences.In informal listening only mixed speaker sentences where the carrier was the male imitating the female were heard as wholly originating from a single speaker.Recordings were made in an lAC 120] booth using an Sennheiser MKH 40 microphone.These sounds were ampli ed (Revox A77), low passed filtered at 9 kHz with a 48-dB per octave cutoff slope (Kemo VBFS).digitised with a 16 bit resolution at a sampling frequency of 20 kHz (Data Translation DTZSZ3) and stored with the ".5 program RDA running on n Victor PC186 computer.Digital wavefomts of the carrier sentences and test words were transferred to a Sun Sparcstation computer for processing.The vowels from the test words were analyzed to find the rst and second formants.Fm were made of harming windowed seynents of each vowel, and a lattice linear prediction (LP) analysis was used to approximate the vocal tract response.For the male speaker the LP analysis was obtained from a 15-ms frame and for the femalc

Read the paper · More papers on PaperTik