Efficient Convolutional Patch Networks for Scene Understanding
Clemens-Alexander Brust, Sven Sickert, Marcel Simon, Erik Rodner, Joachim Denzler · 2015
In this paper, we present convolutional patch networks, which are convolutional (neural) networks (CNN) learned to distinguish different image patches and which can be used for pixel-wise labeling. We show how to easily learn spatial priors for certain categories jointly with their appearance. Experiments for urban scene understanding demonstrate state-of-the-art results on the LabelMeFacade dataset. Our approach is implemented as a new CNN framework especially designed for semantic segmentation with fully-convolutional architectures. In the last years, the revival of convolutional (neural) networks (CNN) [5] has led to a breakthrough in computer vision and visual recognition. While the majority of works focuses on applying these techniques for object classification tasks, there is another field where CNNs can be really useful: semantic segmentation, i.e., assigning a class label to each pixel in an image. In this paper, we show how to learn spatial priors during CNN training, because some classes appear more frequently in some areas of an image. In general, predicting the label of a single pixel requires a large receptive field to incorporate as much context information as possible. We avoid this by incorporating absolute position information in a layer of the CNN as additional input. Urban scene understanding features a number of categories that need to be distinguished, such as buildings, cars, sidewalks, etc. We obtain state-ofthe-art performance in this domain on the LabelMeFacade dataset [4]. Architecture and CNN training Convolutional (neural) networks (CNNs) [5] are feed forward neural networks, which concatenate several layers of different types with convolutional layers playing a key role. The main idea is that the whole classification pipeline consists of one model, which can be jointly optimized during training. The goal of our network is to predict the object category for every single pixel in an image. The CNN architecture is completely described in [1]. However, in addition to [1], we implemented a fully-convolutional version [6] of it which input image convolutional, pooling and non-linear activation layers label estimation and segmentation absolute patch location Figure 2. Basic outline of our CNN architecture. The x and y feature maps allow for learning a spatial prior. is mathematically equivalent, but allows for fast prediction. We still train the network in a patch-wise manner, since preliminary experiments showed that training the network in a fully-convolutional manner (batches for gradient computation are comprised of full images only) resulted in slower (wall time) convergence and ultimately a less accurate network, in contrast to the results of [6] on other datasets. With image-based gradient batches, the model only learned to distinguish between the four most common classes. This may be due to our relatively small dataset resulting in a reduced randomization during optimization, altough we try to introduce more randomness by using spatial loss sampling as detailed in [6]. Incorporating spatial information Predicting the category by only using the information from a limited local receptive field can be challenging and in some cases impossible. We exploit that the absolute position of certain categories in the image is an important contextual cue. We provide the normalized position of a patch as an additional input to the CNN. In particular, the x ∈ [0, 1] and y ∈ [0, 1] coordinates are added as additional feature maps to one of the layers (Figure 2). Whereas incorporating the position information is a common and simple trick in semantic segmentation, with [4] being only one example, combining these priors with CNN feature learning has not been exploited before. New CNN library: CN24 We implemented a new open source CNN library specifically designed for semantic segmentation [1], which is publicly available. An important