A computational theory for locating human faces in photographs
Venugopal Govindaraju · 1992
The human face is an object that is easily located in complex scenes by infants and adults alike. Yet the development of an automated system to perform this task is extremely challenging. An attempt to solve this problem raises two important issues in object location. First, natural objects such as human faces tend to have boundaries which as yet have not been accurately described by analytical functions. This renders the commonly used parameter-based techniques like the Hough transform inadequate for extracting the shape. Second, the object of interest could occur in a scene in various sizes, thus requiring scale independent techniques which can detect instances of the object at all scales. Although, the task of identifying a well-framed face (as one of a set of labeled faces) has been well researched, the task of locating a face in a natural scene is relatively unexplored. We present a computational theory for locating human faces in scenes with certain constraints. Our experiments will be confined to instances where people's faces are the primary subject of the scene, occlusion is minimal, and the faces contrast well against the background. A hypothesis generate-and-test paradigm is proposed and justified as a methodology for face location. Alternative methods of hypothesis testing by either performing rigorous face-specific analysis or using collateral information have been addressed. The shape of the object is defined in terms of features selected using cognitive principles of human perception. Geometrical relationships between features are not rigid. Rather, they are represented by spring functionals that allow several configurations of the features to match against the model. The framework of spring functionals provides a mathematical basis for evaluating the goodness of matches between the data and the model.