Traitement automatique et analyse de la variation dans la parole : des mesures phonétiques sur grands corpus aux réseaux de neurones profonds
Cédric Gendrot · HAL (Le Centre pour la Communication Scientifique Directe) · 2021
This document goes over my teaching, administrative, and research activities since my recruitment as an associate professor at the University Sorbonne Nouvelle in 2006. This summary focuses on the last point, following the common thread of my work: the use of large corpora of unprepared speech for automatic phonetic analysis in order to better understand the variation present in speech.In the first section, after providing formants reference values for French, I showed acoustic reduction phenomena for all vowels as a function of their phonetic duration, consonantal context and speech style. This reduction is also observed in several languages with different phonological constraints. It has been shown in the course of this work that measurements performed automatically on automatically aligned corpora remain consistent provided that certain methodological precautions are respected. In the second section, I highlighted the importance of prosody on the acoustic realization of vowels. Word position, accentual phrase, and intonational phrase are three recurring factors of variation found in French, German, and Spanish. The comparison between three languages with different accentual systems allowed me to separate the accentual structure and the prosodic structure, which can be enhanced respectively either by spectral information (formants) in a preponderant way, or by prosodic parameters (f0 and duration). In the third section, I dealt with linguistic phenomena whose variation raises phonological issues. I was able to show in the context of schwa analysis that taking into account multiple factors was possible and useful in large corpora. The identification of different variables for the reduction of schwa vs. its complete elision allowed us to conclude that there are different mechanisms, one phonetic and the other phonological. Analysis of standard French /R/ from a combination of articulatory and large speech corpora allowed us to consider the unvoiced form of /R/ as the hyper-articulated realization of the voiced form, and showed that /R/ variation is greatly influenced by prosodic position and speech style, in addition to its consonantal context. Finally, in a study postulating that /e/ and /ɛ/ have entered a process of merging, I showed that large corpora with multiple speakers are perfect tools for spotting global trends in a language despite the maintenance of inter-speaker variation. These studies were also an opportunity to perceptually test the measured variations and thus validate their relevance in the context of spoken communication. Several fundamental methodological aspects and innovative methods are presented.In the fourth and last section, a discussion is proposed: the use of large corpora is compared to that of small corpora of read speech. A reappraisal of the methods, both for data and for analysis, is also presented and solutions are proposed. My recent work has led me to the search for speaker-specific strategies and phonetic characterization. For less than ten years, deep neural networks have revolutionized the field of classification, and it seemed essential to try and use them for phonetic analysis. By using convolutional neural networks (CNNs) through spectrograms, the goal was twofold: (1) to know to what extent the spectrogram can characterize the speaker beyond a classical phonetic analysis and (2) by means of visualization techniques, to localize the areas of the spectrogram used by the CNNs. Encouraging results presented in the final discussion provide a glimpse of my future research projects.