Comparing L1 and L2 phoneme trajectories in a feature space of sound and midsagittal ultrasound tongue images

Keita Sano, Yuichi Yaguchi, Ian Wilson · The Journal of the Acoustical Society of America · 2012

To support the development of pronunciation training systems for non-native (L2) speakers, past research has proposed visualization of a speaker's tongue using ultrasound as feedback showing differences between L2 and native (L1) speakers. However, there has been little or no quantitative assessment combining temporal variation of speech sounds and ultrasound tongue images. We propose a mining method to analyze such temporal differences between L1 and L2 speakers. We firstly construct two eigenspaces: one made from feature vectors of speech sounds using Spectrum Vector Field (SVF) and the other from ultrasound tongue images using Histogram of Oriented Gradients (HOG). Next, we compare the movements of L1 and L2 trajectories. Furthermore, we model the connection of phonemes by finding tongue shapes from adjacent speech sounds, and we indicate the differences between L1 and L2 speakers to make a clear intermediate representation from the feature space. In our experiment, we analyze the differences between L1 and L2 pronunciation by focusing on the temporal trajectories of the feature space. These trajectory differences between L1 and L2 speakers’ speech sounds will be presented. We will also present the feature space of ultrasound tongue images that indicate the intermediate tongue shapes mentioned above.

Read the paper · More papers on PaperTik