A principled approach to remove false alarms by modelling the context of a face detector

Cosmin Atanasoaei, Chris McCool, Sébastien Marcel · 2010

Face detection [1, 6] is the task of classifying a sub-window as being a face or not. There are many ways to obtain sub-windows from an image, with the sliding window approach being the most well known. This can result in multiple detections and false alarms. A merging and pruning heuristic algorithm is then typically used to output the final detections [3]. Recent work has been done to overcome the limitations of the sliding window approach by using a branch-and-bound technique to evaluate all possible sub-windows in an efficient way [2]. A different approach was recently proposed in [4] and [5] where they show that the score distribution is significantly different around a true object location than around a false alarm location. We propose a model to enhance a given face classifier, by discriminating false detections (sub-windows) from true detections using the contextual information. Our approach follows the work of [4, 5], but we propose a more discriminative approach and we extract a larger variety of features. We investigate the detection distribution around some sub-window (which we call the context) from which we compute features from every possible axis combination (location and scale). The main advantages of our method is that it can be initialized with any sub-window collection and it poses no restriction regarding the object classifier to run on top of. To build the context of a target sub-window Tsw = (x,y,s), we sample in the 3D space of location (x,y) and scale (s) to collect detections. Then the context of Tsw consists of collection of 4D points C(Tsw) = {(xi,yi,si,msi)i=1,..}, where ms is the classifier score. We propose two strategies for context sampling: full and axis. The full strategy consists of sampling by varying the location and scale at the same time, while the axis strategy the sampling is done just along one axis at a time. The feature vectors are defined by their attribute and the axis combination (x, y and s) used to obtain the attribute. We use 5 attributes that capture the global information (counts), the geometry of the detection distribution (hits) and the detection confidence (score) obtained from the face classifier.

Read the paper · More papers on PaperTik