Constrained optimization for a speech driven talking head
Kyoung-Ho Choi, Jonghoon Lee · 2003
In this paper, a novel algorithm for audio-to-visual conversion based on constrained optimization is presented. Based on facial muscle analysis, the dynamics of mouth movements are modeled and constraints are obtained from them. The obtained constraints are used to estimate visual parameters from speech in a framework of HMM-based visual parameter estimation. The proposed constrained optimization approach finds visual parameters that satisfy given constraints and maximize the auxiliary functions that are used to train the audio-visual HMMs. This approach enables the algorithm to produce reliable visual parameters even in noisy environments. Experimental results demonstrate that the proposed audio-to-visual conversion method is able to follow true visual parameters robustly in various noisy environments.