Verification of Articulatory Phonetics Features with Quantitative Data
László Czap · Acta Polytechnica Hungarica · 2023
This paper aims to refine the base data set of visemes -the visual counterparts of phonemes -with quantitative data to provide accurate input for visual speech synthesis (a talking head that supports the training of speech production of deaf and hard-of-hearing children).Measurement-based features extend the existing data and refine our previously used dynamic model of articulation.This requires the definition of two major types of data simultaneously: the shape of the mouth, which can be examined relatively simply in an ordinary camera image, and the position of the tongue, the analysis of which requires the use of medical-level imaging devices and the processing of their signals.Articulatory phonetics can be divided up into three areas to describe consonants.These are voice, place, and manner respectively.This study aims to confirm the description of the place of articulation with measurement data.Data derived from the shape and position of the tongue is suitable for determining the place of articulation of sounds.In the case of vowels, we estimated the tongue position with the centroid of the tongue while in the case of consonants, we define the place of articulation with the measured distance of the tongue from the palate.To measure these, we used MRI and US images and determined tongue contours with an automated process.The results of this analysis statically define data for articulation keyframes for visual speech synthesis.We applied our results to improve the existing Hungarian transparent talking head with a more accurate model based on the clarification of the dynamic features.We also adapted the same model to the Chinese Shaanxi Xi'an dialect.