Automatic caption generation for video data. Time alignment between caption and acoustic signal

Kazuhiro Watanabe, Masashi Sugiyama · 1999

This paper discusses automatic caption generation, and specifically focuses on correspondence between Japanese text and its speech data. This paper proposes the time alignment module implemented using DP matching and evaluates its performance. Optimizing weight and DP path, the caption display time gap between correct and estimated is less than 39.0 ms in the phoneme boundary. Effects of other speaker's phoneme templates and text phrase deletion are evaluated.

Read the paper · More papers on PaperTik