VizObj2Vec: Contextual Representation Learning for Visual Objects in Video-frames

Ahnaf Farhan, Mahmud Shahriar Hossain · 2020

While the use of the distributional hypothesis has become popular in creating embedding for text corpus, it is rarely used for generating the contextual (distributed) representation of visual objects in video data. In this paper, we present a distributed representation model, vizObj2Vec, that leverages the contexts of visual objects learned from spatiotemporal placements of the objects in the video-frames to construct object-embeddings. The model enables computation of contextual similarity - rather than solely relying on the visual resemblance in similarity computation - between a pair of visual objects. As a result, objects that are contextually connected - because they are in close proximity in a frame or are in nearby frames - appear in a neighborhood in the constructed embedding space. Through a series of extensive experiments, the paper demonstrates (1) the context selection process for visual objects in video data and (2) the potential of the proposed model in distinguishing neighborhoods of contextual objects in the video embedding space.

Read the paper · More papers on PaperTik