3D Vision Technology for Capturing Multimodal Corpora: Chances and Challenges

Gabriele Fanelli, Jüergen Gall, Harald Romsdorfer, Thibaut Weise, Luc Van Gool · 2010

Data annotation is the most labor-intensive part for the acquisition of a multimodal corpus. 3D vision technology can ease the annotation process, especially when continuous surface deformations need to be extracted accurately and consistently over time. In this paper, we give an example use of such technology, namely the acquisition of an audio-visual corpus comprising detailed dynamic face geometry, transcription of the corpus text into the phonological representation, accurate phone segmentation, fundamental frequency extraction, and signal intensity estimation of the speech signals. By means of the example, we will discuss the advantages and challenges of integrating non-invasive 3D vision capture techniques into a setup for recording multimodal data. 1.

Read the paper · More papers on PaperTik