Methodologies for Modelling Silent Pause Length. Insights on Individual and Situational Variation from a Large-Scale Corpus Study.

George Christodoulides, Iulia Grosman, Liesbeth Degand, Anne-Catherine Simon · Digital Access to Libraries (Université catholique de Louvain (UCL), l'Université de Namur (UNamur) and the Université Saint-Louis (USL-B)) · 2017

The temporal organisation of speech can be studied by segmenting the speech signal into measurable components, as a sequence of articulated intervals and pauses. Silent pauses fulfil multiple functions (Zellner, 1994), ranging from the most basic (e.g. breathing, or pre-occlusive pauses linked to articulation), to prosodic functions (e.g. as an important correlate of prosodic phrasing, cf. Krivokapic, 2007; Simon & Christodoulides, 2016), and discursive/rhetorical functions (e.g. to indicate saliency and focus, cf. Duez, 1982). The statistical analysis of silent and filled pause length presents a number of methodological challenges, as the typical distribution of pause durations is positively skewed; therefore the use of methods that rest upon the hypothesis of normality is not appropriate (Oehmen, 2010). An alternative method is to study the distribution of the logarithm of pause durations. For example, Kirsner & Hird (2005) find a bimodal distribution of the log-transformed silent pause lengths in their corpus and posit that the first component distribution (short pauses) corresponds to articulatory processes, while the second component (medium-length pauses) corresponds to cognitive processes, including discourse segmentation. However, the method of log-transformation must be applied with caution, after establishing that the original pause length distribution is indeed bimodal. Campione & Véronis (2002) analysed 5 hours of read and spontaneous speech in five languages, and report a tri-modal distribution of pause length, categorizing them as brief (less than 200 ms), medium (200 to 1000 ms) and long (over 1000 ms). They only found long (>1s) pauses in spontaneous speech, and reported that pauses follow a log-normal distribution globally and for each category. Demol et al. (2007) analysed a 4-hour corpus of three different speaking styles, in six European languages, finding that the “logarithmic duration of the pauses can be well approximated by a bi-Gaussian distribution” both in slow and in fast speaking rates; similar pausing strategies were found for all languages (Dutch, English, French, Italian, Romanian and Spanish). Goldman et al. (2010) studied a 40-minute French spoken corpus with 4 speaking styles (reading, narration, broadcast news and university lectures): they report a multimodal distribution of log-transformed pause length and propose to model silent pauses as a mixture of log-normal distributions. Furthermore, research in perception and psycholinguistics has shown that perceived pauses do not correspond to physical pauses. This is a manifestation of a more general property of the human sensory system: the perception threshold is higher than the actual physical stimulus (Zellner, 1994: 43), and is modulated by previously presented stimuli. We could therefore envisage modelling pausing behaviour using a measure that normalises the duration of each silent pause based on local speech rate, including the length of silent pauses in its immediate context: i.e. using relative length values, rather than absolute length values, or the logarithm of absolute length values. We seek to address these methodological questions through a large-scale corpus study and statistical analysis of the properties of different measures of silent pause length. We have compiled five phonetically aligned corpora of French speech: the LOCAS-F corpus (Martin, Degand, and Simon 2014), the C-Humour corpus (Grosman 2016), the Driving Simulator Cognitive Load corpus (Christodoulides 2016), the C-Phonogenre corpus (Prsir, Goldman, and Auchlin 2014), and the Rhapsodie corpus (Lacheret et al. 2014). The compilation covers 31 speaking styles, includes a total of 276 different samples, and its total duration is 17.4 hours (186.895 tokens). As the corpus contains both monologues and dialogues, we are only focusing on pauses inside a speaker’s turn (between-speaker gaps have been excluded from the analysis). The corpus compilation contains approximately 23.000 turn-internal silent pauses. Figure 1 shows the distribution of four different measures of pause duration: absolute length; the base-10 logarithm of absolute length; the relative duration of each silent pause; and the base-10 logarithm of the relative duration of each silent pause. Relative duration is defined as the length of a pause divided by the arithmetic mean of the length of neighbouring segments within a window of ±5 segments (including both syllables and pauses). As expected, none of the four measures follows a normal distribution. While panel B (log­-transformed durations) may suggest that this measure produces a bimodal distribution (and described as a mixture of Gaussian models), we have applied these methods to each speaking style separately, and found that some speaking styles follow a unimodal distribution, others follow a bimodal distribution, and a few follow a trimodal distribution. Figure 2 shows the distribution of the log­-transformed durations for four different speaking styles (academic speech, radio news, reading and stand-up comedy). We observe that, while log10(duration) appears to follow a bimodal distribution at the speaking style level (i.e. jointly modelling the distribution of pauses from several different speakers), this bimodality disappears at the individual speaker level. As can be seen in Figure 3, individual variation is greater in some speaking styles. Furthermore, the log10(relative duration) measure follows a unimodal distribution in almost all cases. These findings lead us to question the appropriateness of modelling silent pause length as a mixture of log-normal distributions. They also suggest that the local speech rate context may be playing a more important role than previously assumed (a hypothesis that should be tested through targeted perception experiments). Additional analyses will be presented, regarding the relationship between the statistical distribution of silent pause length, and (a) their position in the syntactical structure, as well as (b) their occurrence as part of a disfluency interregnum. In this presentation we focus on methodological questions, on the basis of a corpus study. Perspectives for future research include applying these methods to describe the distribution of filled pause length. This corpus study will facilitate forthcoming perceptual experiments in order to validate which of the modelling methods is closer to the perception of silent pause length.

Read the paper · More papers on PaperTik