Viewing angle estimation for multi-pose AVASR system based on PCNN
Mengjun Wang, Xiangling Wang, Gang Li · 2012
In traditional multi-views audio-visual automatic speech recognition (AVASR) system, Projecting processing is adopted to projecting the features into a uniform pose, which will bring more computation. So different from the past investigations, a different research method is adopted. In this method, different views lipreading used different lip vectors; viewing angle estimation is before feature extraction. Pulse Coupled Neural Network (PCNN) is used to extract features in the gray image sequences of visual speech to estimate the viewing angle. Time series, Entropy series, Logarithm series, and Standard deviation are considered as the feature vector. Experiments are carried out based on Mean Square Error (MSE) in a small database for speaker-dependent case. Experiment results show that feature vector based on PCNN can estimate the viewing angle: 0°, 45°, and 90°. The maximum rate of accurate classification can be reached 95.64% based on Logarithm series.