Tracking and detecting objects in image sequence

Li Wang · 2016

Tracking and detecting arbitrary objects are important in many applications such as video surveillance, video analytics and human-machine interactions.Although many promising methods have been proposed in this area, it is still very challenging to track and detect arbitrary objects due to issues such as complicated motion transformations and occlusions.In this thesis, four pieces of works are developed to address these problems in tracking and detecting objects.The first piece of work addresses learning hierarchical features for visual object tracking by using deep learning.Previously, raw pixel values or hand-crafted features are used to represent target objects.However, these representations are not able to handle large appearance variations of arbitrary target objects.Recently, deep learning has achieved very promising results in speech recognition and image classification.Nevertheless, it is non-trivial to apply deep learning to visual object tracking.Usually, deep neural networks require a lot of training data to learn a large number of network parameters.However, training data is not sufficient for visual object tracking as annotations of a target object are only available in the first frame of a test sequence.To solve this problem, a feature learning algorithm is proposed for visual object tracking by using domain adaptation.First, hierarchical features are learned from auxiliary video sequences by using a two-layer neural network.Embedding the temporal slowness constraint into the stacked network architecture makes learned features robust to complicated motion transformations, and this manipulation is important for visual object tracking.Then, given a target image sequence, a domain adaptation module is proposed to adapt the pre-learned features according to the specific target object.The adaptation is conducted in both two layers of the neural network to include appearance information of the specific target object.As a result, learned hierarchical features can

Read the paper · More papers on PaperTik