Learning discriminative representations from RGB-D video data
Li Liu, Ling Shao · 2013
Recently, the low-cost Microsoft Kinect sensor, which can capture real-time high-resolution RG-B and depth visual information, has attracted in-creasing attentions for a wide range of application-s in computer vision. Existing techniques extract hand-tuned features from the RGB and the depth data separately and heuristically fuse them, which would not fully exploit the complementarity of both data sources. In this paper, we introduce an adap-tive learning methodology to automatically extract (holistic) spatio-temporal features, simultaneously fusing the RGB and depth information, from RGB-D video data for visual recognition tasks. We ad-dress this as an optimization problem using our proposed restricted graph-based genetic program-ming (RGGP) approach, in which a group of prim-itive 3D operators are first randomly assembled as graph-based combinations and then evolved gener-ation by generation by evaluating on a set of RGB-D video samples. Finally the best-performed com-bination is selected as the (near-)optimal represen-tation for a pre-defined task. The proposed method is systematically evaluated on a new hand gesture dataset, SKIG, that we col-lected ourselves and the public MSRDailyActivi-ty3D dataset, respectively. Extensive experimental results show that our approach leads to significan-t advantages compared with state-of-the-art hand-crafted and machine-learned features. 1