On the analysis of speech rhythm for language and speaker identification

Athanasios Lykartsis · DepositOnce · 2020

In the context of this dissertation, novel methods for rhythm description and extraction originating from the area of Music Information Retrieval (MIR) were adapted and applied to represent speech rhythm and its properties. These methods were then used to extract rhythmic information to be used in two specific classification scenarios relevant to speech technology: language identification (LID) and speaker identification (SID). Specifically, periodicity representations that offer an overview of the prominent “beats” – i.e., the salient, recurring temporal or spectral patterns in the audio signal – were created by using the Beat Histogram, an established method for extraction of rhythm information in MIR. The adaptation entailed the analysis of several signal features (e.g., fundamental frequency, energy, spectral change and others) which describe relevant signal properties and directly shape human percepts of, for instance, syllables, phones, accents and prosody. This approach was then thoroughly tested on two multilingual speech datasets with different properties (read vs. spontaneous speech, high vs. low audio signal quality, Indo-European languages only vs. others) using state-of-the- art machine learning algorithms. The results of the experiments for LID show that speech rhythm description based on the proposed methods can be successful, but mostly in the case of read speech with high audio signal quality and for Indo-European languages, pointing towards a potential for improvement of the descriptor robustness. The results are promising, and they surpass the state-of-the-art results of other studies on LID for the used datasets, demonstrating that the proposed features indeed capture a significant part of the variability between languages. Further experiments performed on a dataset of Swiss German for SID showed that rhythmic information is less informative for that task, and that spectral information accounted for much of the variability between speakers. Finally, a feature selection procedure showed descriptors such as tempo (i.e., speech rate), spectral change and fundamental frequency to consistently be among the most useful and informative ones. This finding highlights the fact that it is important to reliably extract salient temporal information, as the descriptors resulting from it are, in many cases, informative as well. Similar results were obtained when the methods were applied to the related task of rhythm-based genre classification on music datasets, suggesting that the findings are not strictly speech specific. Furthermore, listening test experiments for differences in listening to speech vs. listening to singing confirm the findings about the most salient features to be tempo and regularity. Finally, the language rhythm family hypothesis (for example, English and German as the “morse-code” family and Spanish, Italian and French as the “machine-gun” family) could be partially confirmed, but not in its original form. This possibly shows that rhythm classes, which have been difficult to identify with other methods (e.g., other speech rhythm metrics) are also hard to be detected using automatic methods. Alternatively, this might hint at a gap between human- perceived cues and machine-extracted descriptors for speech rhythm. The developed analysis system in the context of the dissertation can be used for rhythm description for various tasks.

Read the paper · More papers on PaperTik