Does one size really fit all?

Christian Hentschel, Harald Sack · 2014

Bag-of-Visual-Words (BoVW) features that quantize and count local gradient distributions in images similar to counting words in texts have proven to be powerful image representations. In combination with supervised machine learning approaches, models for various visual concepts can be learned. While kernel-based Support Vector Machines have emerged as a de facto standard an extensive comparison of different supervised machine learning approaches has not been performed so far. In this paper we compare and discuss the performance of eight different classification models to be applied in BoVW approaches for image classification: Naïve Bayes, Logistic Regression, k-nearest neighbors, Random Forests, AdaBoost and linear Support Vector Machines (SVM) as well as generalized Gaussian kernel SVMs. Our results show that despite kernel-based SVMs performing best on the official Caltech-101 dataset, ensemble methods fall only shortly behind. In addition we present an approach for intuitive heat map-like visualization of the obtained models that help to better understand the reasons of a specific classification result.

Read the paper · More papers on PaperTik