Comparing Visual Feature Coding for Learning Disjoint Camera Dependencies
Xiatian Zhu, Shaogang Gong, Chen Change Loy · 2012
Problem: This work systematically investigates the effectiveness of various visual feature coding schemes for facilitating the learning of timedelayed dependencies among disjoint multi-camera views. Related work: Quite a few studies [3, 4, 6] have been proposed to model inter-camera dependency across non-overlapping camera views. Learning time-delayed correlations among disjoint cameras in crowded public scenarios is a non-trivial task: (1) the time gaps between camera views are unknown therefore activities in two related views may occur at arbitrary time delays with high uncertainty; (2) the features are inevitably noisy, ambiguous, and may vary drastically across views owning to illumination condition, camera angles, and changes in object pose. Most state-ofthe-art methods typically hand pick a few features tailored to the target environment, with the hope that those chosen features contain robust and sufficient statistics for correlating the time-delayed activity patterns across disjoint views. These manual approaches to hard selection of features are neither principled nor generalisable to different scene context. Our solution: In this study, we wish to examine the concept that visual features should be coded and selected automatically for robust and accurate time-delayed dependency learning. The contributions of this study are two-fold: (1) We present a systematic study and evaluation to investigate the effectiveness of supervised and unsupervised feature coding methods to facilitate the learning of inter-camera activity pattern dependencies. (2) We systematically evaluate the sensitivity of inter-camera time delayed dependency learning given different training video sizes and region decomposition qualities. These factors are critical for accurate dependency learning but have been largely ignored by the published existing work in the literature. Approach overview: We employ the Random Forest [2] as the supervised feature coding approach. In particular, given a set of localised features extracted from a region, together with people count training label over time, we first train a regression forest to learn the non-linear mapping between the crowd density and the corresponding low-level features. Given unseen data, we then construct a time series based on the predicted crowd density ŷ obtained from the regression forest (RF pred), the treestructured code (tree code) [5], or the combination of the two. As for unsupervised coding scheme, we use the Latent Dirichlet Allocation (LDA) [1] to map the low-level features into codewords that capture the topic distribution, whereby an image region patch (document) d is treated as a collection of j = 1 . . .Ni features (words). To form the unsupervised feature codes, given a sequence of localised feature vectors detected from a region, we first perform quantisation on each feature to generate a bag-of-word representation for all image patches. Similar to text documents, these bag-of-word represented image patches are fed into the LDA, which gives us a topic-based representation. Once having the topic-based code (topic code), we perform k-means quantisation on them, producing the final compact topic-based code, and concatenate them over time to form a time series. To solve the problem of using the feature codes for learning intercamera dependencies, we adopt the Time Delayed Mutual Information (TDMI) proposed in [3] due to its reported effectiveness and simplicity. The input to TDMI are time series generated from either the supervised or the unsupervised coding scheme. In addition to measuring deviation error in transition time, we propose a new metrics to evaluate the effectiveness of different coding methods, called Mutual Information Margin (MIM):