Publication Details
Overview
 
 
Dongmei Jiang, Yong Zhao, Hichem Sahli, Yanning Zhang
 

Contribution to journal

Abstract 

This paper presents a photo realistic facial animation synthesis approach based on an audio visual articulatory dynamic Bayesian network model (AF\_AVDBN), in which the maximum asynchronies between the articulatory features, such as lips, tongue and glottis/-velum, can be controlled. Perceptual Linear Prediction (PLP) features from audio speech, as well as active appearance model (AAM) features from face images of an audio visual continuous speech database, are adopted to train the AF\_AVDBN model parameters. Based on the trained model, given an input audio speech, the optimal AAM visual features are estimated via a maximum likelihood estimation (MLE) criterion, which are then used to construct face images for the animation. In our experiments, facial animations are synthesized for 20 continuous audio speech sentences, using the proposed AF\_AVDBN model, as well as the state-of-art methods, being the audio visual state synchronous DBN model (SS\_DBN) implementing a multi-stream Hidden Markov Model, and the state asynchronous DBN model (SA\_DBN). Objective evaluations on the learned AAM features show that much more accurate visual features can be learned from the AF\_AVDBN model. Subjective evaluations show that the synthesized facial animations using AF\_AVDBN are better than those using the state based SA\_DBN and SS\_DBN models, in the overall naturalness and matching accuracy of the mouth movements to the speech content.

Reference