An audio-visual speech database and automatic measurements of visual speech

Tobias Öhman · 2007

A database of video sequences of a male speaker uttering 269 Swedish sentences and 153 VCV-words has been recorded. Parts of the speaker's face were marked to facilitate optical measurements. Algorithms to automatically determine the shape of the lips as well as areas and other visual features of the speaker's face have been developed. The audio signal has been phonetically segmented and labelled. This material differs from most other audio-visual databases mainly in two aspects: firstly the recorded material contains naturally reduced and coarticulated continuously spoken sentences, and secondly, a specially constructed device provides the possibility to easily trace jaw movements. Introduction This paper describes the recording and evaluation of a database, together with algorithms for detailed and precise measurements of visual speech. Statistics and analysis of the measurements will be presented in forthcoming papers and will not be reported here. The term visual speech refers to...

Read the paper · More papers on PaperTik