Video Representation with Dynamic Features from Multi-Frame Frame- Difference Images
Michelle Lee, Alexander W. Lee, D. Lee, Soo-Young Lee · 2007
The extraction of dynamic motion features are reported from multiple video frames by three unsupervised learning algorithms, i.e., Principal Component Analysis (PCA), Independent Component Analysis (ICA), and Non-negative Matrix Factorization (NMF). Since the human perception of facial motion goes through two different pathways, i.e., the lateral fusifom gyrus for the invariant aspects and the superior temporal sulcus for the changeable aspects of faces, we extracted the dynamic video features from multiple consecutive frames for the latter. Both the original videos and the frame-difference sequences are used for comparison. The required number of multiframe features for the same representation accuracy is almost independent upon the frame length. Therefore, the multiple-frame features are much more efficient for video representation than the single-frame static features. The extracted features are also used for lipreading, and the features from frame-difference sequences demonstrated better recognition rates than those from original videos.