Parallelizing support vector machines for scalable image annotation
Nasullah Khalid Alham · Brunel University Research Archive (BURA) (Brunel University London) · 2011
Machine learning techniques have facilitated image retrieval by automatically classifying and annotating images with keywords.Among them Support Vector Machines (SVMs) are used extensively due to their generalization properties.However, SVM training is notably a computationally intensive process especially when the training dataset is large.In this thesis distributed computing paradigms have been investigated to speed up SVM training, by partitioning a large training dataset into small data chunks and process each chunk in parallel utilizing the resources of a cluster of computers.A resource aware parallel SVM algorithm is introduced for large scale image annotation in parallel using a cluster of computers.A genetic algorithm based load balancing scheme is designed to optimize the performance of the algorithm in heterogeneous computing environments.SVM was initially designed for binary classifications.However, most classification problems arising in domains such as image annotation usually involve more than two classes.