Per-Exemplar Fusion Learning for Video Retrieval and Recounting
Ilseo Kim, Sangmin Oh, A. G. Amitha Perera, Chin‐Hui Lee · 2012
We propose a novel video retrieval framework based on an extension of per-exemplar learning [7]. Each training sample with multiple types of features (e.g., audio and visual) is regarded as an exemplar. For each exemplar, a localized per-exemplar distance function is learned and used to measure the similarity between itself and new test samples. Exemplars associate only with sufficiently similar test data, which accumulate to identify the data to be retrieved. In particular, for every exemplar, relevance of each feature type is discriminatively analyzed and the effect of less informative features is minimized during the fusion-based associations. In addition, we show that our framework can enable a rich set of recounting capabilities where the rationale for each retrieval result can be automatically described to users to aid their interaction with the system. We show that our system provides competitive retrieval accuracy against strong baseline methods, while adding the benefits of recounting.