Fourier shape-frequency words for actions
Bishwajit Sharma, KS Venkatesh, Amitabha Mukerjee · 2011
Actions consist of short shape-motion fragments which recur in a seemingly unique sequence. We propose that these short fragments may constitute a concise vocabulary for actions. Models based on such “words” sometimes use the bag of words paradigm, which ignores sequence information. Also, despite the well-known utility of Fourier and similar features for temporal modelling, Fourier models have not received due attention to model action words until recently. Hence, we employ shape-frequency features as a temporally windowed Fourier transform to capture local motion and shape information. Unsupervised clustering discovers the naturally occurring modes (words) of these features. Each labelled video can thus be represented as a sequence of cluster transitions. Though different actions share common words, we observe that the word sequences are different for different actions, enabling easy discrimination. We evaluate the model on the Weizmann action dataset [1] and achieve 96.7% classification accuracy, and show how it compares to other similar algorithms.