Image annotation through statistical and interactive learning

Ming Dong, Changbo Yang · 2007

Image annotation is a process in which a computer program automatically assigns keywords to an image by learning the correspondence between image regions and keywords. However, keywords are usually associated with images instead of individual regions in the training data set. This poses a major challenge for any learning strategy. In this dissertation, we present a novel framework for region-based image annotation using Multiple-Instance Learning (MIL). MIL is a variation of supervised learning, where the task is to learn a concept given positive and negative bags of instances. In the framework, images are viewed as bags, each of which contains a few instances corresponding to the segmented image regions. We demonstrate the effectiveness of the proposed MIL solution for image annotation problems through a Sequential Point-Wise Diverse Density (SPWDD) algorithm. Moreover, we propose a novel Asymmetrical Support Vector Machine-based MIL algorithm (ASVM-MIL), which extends the conventional Support Vector Machine (SVM) by introducing asymmetrical loss functions for false positives and false negatives. In turn, this induces a decision boundary that is more distant from one class than from the other. Our experiments show that ASVM-MIL can improve the image annotation performance, especially for keywords with diverse appearance. The experiment results also show that ASVM-MIL can achieve very competitively result on the benchmark MUSK data sets. In addition, we propose an interactive image annotation scheme, in which image annotation is initialized based on statistical modelling and refined through users' interaction and natural language processing. Particularly, we propose a novel relevance feedback solution: semantic feedback, which allows an annotation system to interact with users directly at the semantic level. The learning mechanism built in the semantic feedback can substantially improve the image annotation performance in the long-run.

Read the paper · More papers on PaperTik