Knowledge transfer in object recognition

Lan Xu · Queen Mary Research Online (Queen Mary University of London) · 2020

Object recognition is a fundamental and long-standing problem in computer vision.Since the latest resurgence of deep learning, thousands of techniques have been proposed and brought to commercial products to facilitate people's daily life.Although remarkable achievements in object recognition have been witnessed, existing machine learning approaches remain far away from human vision system, especially in learning new concepts and Knowledge Transfer (KT) across scenarios.One main reason is that current learning approaches address isolated tasks by independently training predefined models, without considering any knowledge learned from previous tasks or models.In contrast, humans have an inherent ability to transfer the knowledge acquired from earlier tasks or people to new scenarios.Therefore, to scaling object recognition in realistic deployment, effective KT schemes are required.This thesis studies several aspects of KT for scaling object recognition systems.Specifically, to facilitate the KT process, several mechanisms on fine-grained and coarse-grained object recognition tasks are analyzed and studied, including 1) cross-class KT on person re-identification (reid); 2) cross-domain KT on person re-identification; 3) cross-model KT on image classification; 4) cross-task KT on image classification.In summary, four types of knowledge transfer schemes are discussed as follows:Chapter 3 Cross-class KT in person re-identification, one of representative fine-grained object recognition tasks, is firstly investigated.The nature of person identity classes for person re-id are totally disjoint between training and testing (a zero-shot learning problem), resulting in the highly demand of cross-class KT.To solve that, existing person re-id approaches aim to derive a feature representation for pairwise similarity based matching and ranking, which is able to generalise to test.However, current person re-id methods assume the provision of accurately cropped person bounding boxes and each of them is in the same resolution, ignoring the impact of the background noise and variant scale of images to cross-class KT.This is more severed in practice when person bounding boxes must be detected automatically given a very large number of images and/or videos (un-constrained scene images) processed.To address these challenges, this chapter provides two novel approaches, aiming to promote cross-class KT and boost re-id performance.1) This chapter alleviates inaccurate person bounding box by developing a joint learning deep model that optimises person re-id attention selection within any auto-detected person bounding boxes by reinforcement learning of background clutter minimisation.Specifically, this chapter formulates a novel unified re-id architecture called Identity DiscriminativE Attention reinforcement Learning (IDEAL) to accurately select re-id attention in auto-detected bounding boxes for optimising re-id performance.2) This chapter addresses multi-scale problem by proposing a Cross-Level Semantic Alignment (CLSA) deep learning approach capable of learning more discriminative identity feature representations in a unified end-to-end model.This provide any help once it is benefical to my research and daily life.The brainstorm in front of the office's whiteboard and the time of we burn the midnight oil to catch up the deadline, are fond memories of my life.Meanwhile, I convey my special thanks to Dr. Xiatian Zhu for his excellent guidance and invaluable

Read the paper · More papers on PaperTik