Deep Spatial Pyramid Ensemble for Cultural Event Recognition
Xiu-Shen Wei, Bin-Bin Gao, Jianxin Wu · 2015
Semantic event recognition based only on image-based cues is a challenging problem in computer vision. In or-der to capture rich information and exploit important cues like human poses, human garments and scene categories, we propose the Deep Spatial Pyramid Ensemble framework, which is mainly based on our previous work, i.e., Deep Spa-tial Pyramid (DSP). DSP could build universal and power-ful image representations from CNN models. Specifically, we employ five deep networks trained on different data sources to extract five corresponding DSP representations for event recognition images. For combining the comple-mentary information from different DSP representations, we ensemble these features by both “early fusion ” and “late fusion”. Finally, based on the proposed framework, we come up with a solution for the track of the Cultural Event Recognition competition at the ChaLearn Looking at Peo-ple (LAP) challenge in association with ICCV 2015. Our framework achieved one of the best cultural event recogni-tion performance in this challenge. 1.