Continuous Gesture Recognition through Selective Temporal Fusion

Pradyumna Narayana, J. Ross Beveridge, Bruce A. Draper · 2019

Gesture recognition is an important task with the potential to revolutionize human/computer interfaces (HCI). Gestures, however, are dynamic. While a few gestures may be static poses, most gestures are complex temporal sequences of motions. For most HCI applications, gestures must be recognized in real-time in streaming data. Therefore, most recognition systems analyze each frame as it comes in, fusing data across time to detect gestures. This paper presents results of the first systematic study of temporal fusion techniques for streaming gesture recognition. These results show that the choice of the best fusion strategy depends on whether the input is global (i.e. full-frame) or a spatially focused window, and on whether the input is unprocessed RGB or depth depth versus a flow field.This conclusion is then used to extend a state-of-the-art architecture for isolated gesture recognition, FOANet [24], to continuous gesture recognition. The result is a system that established a new state-of-the-art for recognition performance on the ChaLearn ConGD data set, with a mean Jaccard Index of 0.77 compared to the previous best result of 0.61. This paper also establishes a baseline of performance for the newer, continuous version of the NVIDIA dataset.

Read the paper · More papers on PaperTik