Low-level fusion of color, texture and depth for robust road scene understanding

Timo Scharwächter, Uwe Franke · 2015

We propose a novel approach to pixel-level semantic labeling, which aims to rapidly infer the coarse layout of street scenes from color, texture and depth information in a joint fashion using a randomized decision forest. The recovered pixellevel class probability maps provide a general purpose basis to guide more elaborate vision algorithms. To demonstrate the richness of our labeling, we extend the well-known Stixel model to use the semantic labels as input cues. In addition, we employ our generated low-level information as an attention mechanism for a vehicle detector. In both cases, recognition performance and accuracy are significantly improved. In our experimental evaluation on the public KITTI benchmark, we thoroughly study the characteristics of different feature channels as well as their contribution to the overall pixellevel labeling result. Our results underline that the combination of several orthogonal feature channels in a joint model is key to superior performance. This performance improvement comes at little additional cost, given that our approach is able to operate at 100 Hz using a GPU implementation.

Read the paper · More papers on PaperTik