Video-based Face Recognition using Spatio-Temporal Representations
John See, Chikkannan Eswaran, Mohammad Faizal Ahmad Fauzi · InTech eBooks · 2011
Face recognition has seen tremendous interest and development in pattern recognition and biometrics research in the past few decades. A wide variety of established methods proposed through the years have become core algorithms in the area of face recognition today, and they have been proven successful in achieving good recognition rates primarily in still image-based scenarios (Zhao et al., 2003). However, these conventional methods tend to perform less effectively under uncontrolled environments where significant face variability in the form of complex face poses, 3-D head orientations and various facial expressions are inevitable circumstances. In recent years, the rapid advancement in media technology has presented image data in the form of videos or video sequences, which can be simply viewed as a temporally ordered collection of images. This abundance and ubiquitous nature of video data has presented a new fast-growing area of research in video-based face recognition (VFR). Recent psychological and neural studies (O’Toole et al., 2002) have shown that facial movement supports the face recognition process. Facial dynamic information is found to contribute greatly to recognition under degraded viewing conditionsand alsowhen a viewer’s experience with the same face increases. Biologically, the media temporal cortex of a human brain performs motion processing, which aids the recognition of dynamic facial signatures. Inspired by these findings, researchers in computer vision and pattern recognition have attempted to improve machine recognition of faces by utilizing video sequences, where temporal dynamics is an inherent property. In VFR, temporal dynamics can be exploited in various ways within the recognition process (Zhou, 2004). Some methods focused on directly modeling temporal dynamics by learning transitions between different face appearances in a video. In this case, the sequential ordering of face images is essential for characterizing temporal continuity. While its elegance in modeling dynamic facial motion signatures and its feasibility for simultaneous tracking and recognition are obvious, classification can be unstable under real-world conditions where demanding face variations can caused over-generalization of the transition models learned earlier. In a more general scenario, VFR can also be performed by means of image sets, comprising of independent unordered frames of a video sequence. A majority of these methods characterize the face manifold of a video using two different representations – (1) face subspaces1, and (2)