Text-driven automatic image sequence generation using facial modeling for digital TV news production system
Terence Chun-Ho Cheung, Hon-Wah Wong, Lai-Man Po · 2002
This paper presents a facial modeling approach for automating the head-and-shoulder image sequence generation for digital TV news video clips production which is the most expensive part in terms of manpower and cost. With the IPA phonetics transcribed from news script being the driving parameter, the high precision adapted 2-D wireframe model on a frontal view of speaker image with sufficient facial textural information will be used to define the associated facial action units (AFAUs) corresponding to the phonemes. This developed facial modeling approach increases the intelligibility of facial non-verbal communication for potential audiovisual lip-synch application like TV news video clips, video telephony, Story Teller On Demand (STOD), lip-reading for the deaf or the hearing-impaired.