Representation of prosodic structure in speech using nonlinear methods.

Rashid Ansari, Yi Min Dai, Jian Lou, David M. McNeill, Francis K. H. Quek · NSIP · 1999

As part of an e ort to discover and quantify cues in the modalities of speech, gesture, and gaze, we are developing automated methods of extracting the primitives from the raw multimodal data, and representing and organizing them in new and exible ways. An in-depth study of the acoustic-prosodic properties of speech such as pitch, amplitude, duration, and speaking rate, and their relation to other modalities in conveying discourse-level meaning requires the processing of a large amount of the recorded audio signal. In this paper we describe a procedure to process the fundamental frequency (F0) trace in order to (i) obtain an improved estimate, (ii) segment the processed F0 data, and (iii) nd a parametric representation suitable for the multimodal analysis. The goal is to use exible parametric representations whose parameters can be examined for correlation with suitably extracted gesture and gaze features. Examples of application of the procedure for preprocessing, segmentation and representation of the F0 samples are described.

Read the paper · More papers on PaperTik