Audio-visual evaluation and detection of word prominence in a human-machine interaction scenario

Martin Heckmann · 2012

This paper investigates the audio-visual correlates and the de-tection of word prominence. Subjects were interacting with a computer in a small game which created a broad and a nar-row focus condition. Audio-visual recordings with a distant mi-crophone and without visual markers were made. As acoustic features duration, intensity, fundamental frequency and spectral emphasis were calculated. From the visual channel head move-ments and image transformation based features from the mouth region were extracted. First the results show that the extracted features are significantly different for the two focus conditions (broad and narrow). Based on classification results it is demon-strated that they can be differentiated without knowledge of the word identity with accuracies of approx. 80%. Furthermore, it is shown that the visual channel by itself yields accuracies no-tably better than chance (approx. 65%) and that a combination of both modalities increases performance to approx. 85%.

Read the paper · More papers on PaperTik